Test Structure

The Subtests Inside an IQ Test
What Each One Is Actually For

An IQ score is not a thing that gets measured. It is a number assembled from a dozen or more small tasks, each built to isolate one narrow ability. Knowing what each task is for is the difference between reading a report and guessing at it.

Illustration of the subtest families that make up a full IQ battery

Quick Answer

Updated August 16, 2026 by Structural. A subtest is a single task type administered as a block, scored on its own scale, and combined with other subtests to produce an index score. Full batteries use between ten and twenty of them because no single task measures general ability well enough on its own.

Direct answer: subtests exist because every task measures its target ability plus a pile of irrelevant things, and the only reliable way to cancel the irrelevant part is to measure the same broad ability with several tasks that share the target and differ in everything else. Two vocabulary-shaped tasks share vocabulary. Two tasks that share nothing but reasoning demand isolate reasoning.

This page walks through each major subtest family, states what narrow ability it targets, and flags the specific confounds each one carries. It also covers why interpreting a single subtest score in isolation is the most common mistake in reading a report, and why the evidence base argues against it.

Why a Battery Uses Many Tasks Instead of One

If general cognitive ability is a single thing, the obvious design would be one excellent test of it. Batteries do the opposite, and the reason is worth understanding because it explains every structural decision that follows.

Any single task measures its target plus everything else it happens to require. A matrix reasoning item measures fluid reasoning, and also visual acuity, attention, comfort with abstract figures, willingness to persist, and whatever the examinee was thinking about beforehand. Those extras are not noise in the statistical sense, they are systematic, and they do not average out across repetitions of the same task. Repeat a matrix item a hundred times and you get a very reliable measure of matrix-item-solving, which is not the same as fluid reasoning.

The fix is to measure the same broad ability with tasks that share the target and differ in their extras. If matrix reasoning and figure weights both require fluid reasoning but demand different perceptual skills, different response formats, and different strategies, then what they have in common is closer to the reasoning itself. Everything they do not share pulls in different directions and partly cancels.

This is not a metaphor for what happens statistically, it is a description of it. Factor analysis extracts what a set of measures share and discards what is specific to each. The whole apparatus of indices and composites is a machine for keeping shared variance and throwing away task-specific variance, which is why composite scores are more reliable and more predictive than the subtests they are built from.

It also explains a fact that surprises people reading their first report: the subtest scores are the least trustworthy numbers on the page, and the composite is the most trustworthy, even though the subtests are what you actually did.

From Item to Subtest to Index to Full Scale

Every modern battery uses the same four-level structure, and the names of the levels are worth keeping straight because reports use them without explanation.

Item

A single question or problem. Scored right or wrong, or occasionally on a partial credit scale. Individually almost meaningless.

What IQ measures
Subtest

A block of items of one type. Raw score converted to a scaled score with mean 10 and standard deviation 3.

Scores and percentiles
Index

Two or more subtests measuring one broad ability. Reported on the familiar scale with mean 100 and standard deviation 15.

Cognitive domains

Above the indices sits the full scale score, built from a designated subset of subtests across all domains. It is the number people mean when they say IQ, and it is the most reliable single figure a battery produces because it pools the most information.

The scaled score convention deserves a note because it confuses almost everybody at first. Subtests use mean 10, standard deviation 3. Indices and the full scale use mean 100, standard deviation 15. A scaled score of 13 is one standard deviation above average, exactly equivalent to an index of 115, but the two numbers look nothing alike. If a report shows a subtest at 13 and an index at 115, those are the same statement made twice in different units.

The conversion from raw to scaled is where norms enter. A raw score of 34 on a vocabulary subtest means nothing by itself. It becomes a scaled score only after comparison with an age-matched reference sample, which is why norms are the foundation everything else rests on and why an instrument without published norms cannot produce an interpretable score at all.

Verbal Comprehension Tasks

Verbal subtests measure crystallised ability, written Gc in the Cattell-Horn-Carroll framework: the depth and accessibility of knowledge acquired through experience and education. They are among the most reliable subtests in any battery and among the most culturally loaded.

Vocabulary. The examinee defines words of increasing difficulty. It is consistently one of the strongest single indicators of general ability in the entire battery, which is counterintuitive until you consider what producing a good definition requires: retrieving the concept, distinguishing it from neighbours, and expressing the distinction precisely. Its weakness is exposure. A word you have never encountered cannot be defined regardless of ability, so the score reflects environment as well as capability.

Similarities. Given two words, the examinee states how they are alike. This targets verbal concept formation, and scoring is graded rather than binary: a superordinate category earns more than a shared property, which earns more than a functional association. That graded structure is what lets a single item discriminate across a wide ability range, and it is also why the subtest requires trained scoring or a well specified rubric.

Information. Factual questions about the world. The most purely knowledge-dependent subtest, and the one most affected by educational opportunity. Its value is that it is highly reliable and relatively insensitive to the transient factors that disturb other subtests, which makes it useful as a stable reference point when the rest of a profile looks disturbed.

Comprehension. Questions about social conventions and practical reasoning. It has the widest range of acceptable answers and the loosest scoring, which makes it the least reliable verbal subtest and the one most likely to be dropped from core composites in current batteries.

The shared confound across this family is language exposure. Verbal subtests measure ability as it has been developed through a specific linguistic environment, and a person tested in a second language will score below their capability by an amount that no statistical adjustment can recover. This is a real limitation rather than a criticism, and it is why verbal and nonverbal indices are reported separately rather than merged.

Fluid Reasoning Tasks

Fluid reasoning subtests measure Gf, the ability to solve novel problems where prior knowledge provides little help. They are the closest thing a battery has to a direct measure of general ability, and the family has changed more than any other in recent revisions.

Matrix Reasoning. A grid of figures with one cell missing, and options to complete it. The examinee must identify the rule governing the grid and apply it. It is the archetypal fluid reasoning task and the one most online tests use exclusively, partly because it is genuinely good and partly because it is easy to generate and score automatically. Its confound is that it is entirely visual, so people who reason well verbally but find abstract figures unnatural are penalised.

Figure Weights. Balance scales with shapes, where the examinee must determine what maintains equilibrium. This targets quantitative reasoning within the fluid domain and has an important property: it is a deductive task with a single defensible answer, so unlike matrix items it cannot be solved by pattern-matching or elimination. That makes it unusually resistant to the guessing strategies that inflate matrix scores.

Picture Concepts. Rows of pictures from which the examinee selects one per row sharing a characteristic. Used mainly in child batteries, and generally the weakest member of the family, because the shared characteristic often depends on knowledge rather than reasoning.

Arithmetic. Word problems solved mentally without paper. Its classification has been argued about for decades, because it demands working memory to hold the problem, quantitative knowledge to know the operations, and fluid reasoning to select them. Different batteries assign it to different indices, which is a straightforward admission that it loads on more than one.

The structural point about this family is that the shift from a single perceptual index to separate fluid reasoning and visual spatial indices was one of the most consequential revisions in modern test design. It separated reasoning about relationships from manipulating shapes, which had previously been merged, and it aligned the batteries with the CHC framework that the research literature had already adopted.

Visual Spatial Tasks

Visual spatial subtests measure Gv, the ability to generate, hold, and transform mental representations of visual material. The family is small because good measures are hard to build.

Block Design. The examinee reproduces a pattern using coloured blocks, against a time limit with bonuses for speed. It is the oldest subtest still in wide use and remains one of the most informative, partly because performance can be observed rather than just scored: how somebody approaches the task often says more than the total. Its confounds are motor control and speed, both of which are timed along with the spatial ability.

Visual Puzzles. A completed shape is shown with pieces below, and the examinee selects those that would assemble it. This is Block Design without the motor component, which is exactly why it was introduced: it isolates mental construction from manual construction. The trade is that it becomes partly a process of elimination for examinees who work systematically.

Spatial rotation and navigation tasks. Less standardised across batteries but common in research instruments and computer administered tests, where rotation can be presented dynamically in ways paper cannot support. These target narrower components of Gv and are the family most improved by computer administration.

Visual spatial ability is the domain with the largest and most consistently replicated average difference between groups, and it is also the domain where practice effects are largest. Both facts argue for interpreting Gv carefully rather than as a fixed characteristic, and neither justifies dropping it, because spatial ability has independent predictive value for outcomes in technical fields that verbal and quantitative measures miss.

Working Memory Tasks

Working memory subtests measure Gwm, the capacity to hold information in an active state while operating on it. The distinction that matters here is between storage and manipulation, and good subtests measure both.

Digit Span. Three conditions in current batteries. Forward repetition is nearly pure storage. Backward repetition requires holding the sequence while reversing it. Sequencing requires reordering into ascending order, which is the most demanding of the three. The gap between forward and backward performance within a person is one of the few subtest-level comparisons with a defensible interpretation, because the two conditions share everything except the manipulation demand.

Letter-Number Sequencing. A mixed string of letters and digits must be reordered, digits ascending then letters alphabetically. It is the most demanding storage-plus-manipulation task in most batteries and correlates well with general ability, but it is also the one most disrupted by anxiety, which limits its usefulness in exactly the situations where working memory is most in question.

Picture Span and visual working memory tasks. Visual analogues of digit span, added to child batteries to avoid making the working memory index entirely dependent on verbal material. Their inclusion improved the construct coverage of the index at some cost in reliability.

Working memory carries a specific interpretive hazard. It is the domain most sensitive to state factors, meaning attention, anxiety, sleep, and medication, and it is therefore the domain where a single low score is least likely to reflect a stable characteristic. A depressed working memory index in a person who was tested while anxious is a finding about that morning, not about that person, which is why clinical interpretation always asks about conditions before drawing conclusions.

Processing Speed Tasks

Processing speed subtests measure Gs, the rate at which simple, overlearned operations are performed under sustained attention. They look the least impressive and carry more clinical weight than their simplicity suggests.

Coding. A key pairs symbols with digits, and the examinee copies matching symbols as fast as possible within a fixed window. Nearly a pure speed measure, with an incidental learning component: examinees who memorise part of the key go faster, which is a small confound that publishers accept because it is common to everybody.

Symbol Search. The examinee decides whether target symbols appear in a search group, repeatedly and against the clock. It emphasises visual scanning and decision speed rather than the motor output that dominates Coding, which is why the two are used together.

Cancellation. Scanning a structured or random array to mark targets while ignoring distractors. Adds a selective attention component and is often supplemental rather than core.

The common confound is motor. All three require a physical response, so fine motor difficulty, an unfamiliar input device, or a hand injury depresses the score without any change in cognitive speed. On computer administration, the pointing device becomes part of the measurement whether the test intends it or not, which is covered in Processing Speed and in Why IQ Tests Are Timed.

Quantitative Tasks

Quantitative subtests measure Gq, acquired mathematical knowledge and the fluency with which it is applied. Not every battery reports Gq as a separate index, and that decision has consequences.

The Wechsler tradition folds quantitative material into other indices, placing Arithmetic under working memory and Figure Weights under fluid reasoning. The Woodcock-Johnson tradition reports quantitative ability separately, treating mathematical achievement as its own broad ability rather than a manifestation of others.

The argument for separating it is that Gq behaves like Gc rather than Gf. It is knowledge that has been acquired, it grows with instruction, and it is not interchangeable with the reasoning that helps you acquire it. Somebody with strong fluid reasoning and weak mathematical instruction has a real quantitative deficit that a battery folding Gq into Gf will not report, because the fluid reasoning items carrying quantitative content are chosen to minimise the knowledge required.

The argument against separating it is contamination. A quantitative index is heavily determined by educational history, which makes it a poor measure of ability in populations with heterogeneous schooling. That is a genuine limitation, and it is the reason clinical batteries that must work across diverse backgrounds have been reluctant to include it as a core index.

The practical position is that a quantitative index is informative when the question is what somebody can currently do, and misleading when the question is what somebody could learn to do. Both are legitimate questions, and a report that does not say which one it is answering is not doing its job.

Core, Supplemental, and Substituted Subtests

Batteries designate some subtests as core, meaning they contribute to the composite scores, and others as supplemental, meaning they are administered for additional information or as replacements. The distinction is not about quality and it is frequently misread.

A supplemental subtest is often excellent. It may be supplemental because the index already has enough subtests, because it adds administration time out of proportion to what it contributes, or because it duplicates coverage. Figure Weights was supplemental in one battery generation and core in the next, without the subtest itself changing.

Substitution rules allow a supplemental subtest to replace a core one when the core subtest was spoiled, for instance by an interruption or an administration error. The rules are restrictive: usually only one substitution per index and a limited number per full scale, and the substitute must come from the same index.

The reason for the restriction is that norms were built on a specific combination. Every substitution moves the actual test slightly away from the instrument the norms describe, and the error introduced is not reported anywhere on the score sheet. One substitution is a small, documented compromise. Several stop being a compromise and start being a different test, which is why manuals set hard limits rather than leaving it to judgment.

This is also why a report should state which subtests were administered and whether any were substituted. A full scale score with two substitutions is a weaker number than the same figure obtained from the standard combination, and nothing in the number itself reveals the difference.

Why Single Subtest Scores Mislead

The most common error in reading a report is treating a single low or high subtest score as a finding. It is usually not, and the reasons are technical rather than a matter of caution.

The first is reliability. Subtests are short, typically fifteen to thirty items after discontinue rules take effect, and short measures are unreliable. A subtest with a reliability coefficient around 0.85 has a confidence interval spanning several scaled score points, meaning a score of 8 and a score of 11 may not be distinguishable at all. Composites pool across subtests and are correspondingly tighter.

The second is that subtest specificity is small. When you partition a subtest's variance into the part shared with general ability, the part shared with its index, and the part unique to that subtest, the unique portion is usually the smallest and is largely error. There is often not enough reliable subtest-specific variance to support an interpretation about that subtest in particular, which is the finding that motivated the shift away from profile analysis in clinical practice.

The third is multiplicity. A battery with sixteen subtests produces one hundred and twenty pairwise comparisons. If you look for the largest gap you will always find one, and in a normal profile the expected largest gap is substantial. Reporting it as a discovered weakness is finding a pattern in noise, and it is the mechanism behind most of the strengths-and-weaknesses narratives that circulate about test profiles.

What survives all three is the index level and above. Index scores pool several subtests, have tighter confidence intervals, and correspond to constructs with independent evidence for their existence. A twenty point gap between two indices is worth discussing. A three point gap between two subtests is not, and the distinction is the single most useful thing to know when reading your own report.

There is a narrow exception. Comparisons designed into the instrument, such as digit span forward against backward, are legitimate because the two conditions were built to differ in exactly one respect and the manual supplies base rates for the difference. That is a planned contrast with a reference distribution, not a gap discovered by scanning.

The base rate point is worth expanding, because it is the part most often left out. Manuals publish how frequently a given difference occurs in the standardisation sample, and the numbers are consistently larger than intuition suggests. Differences that feel dramatic when you see them on your own report turn out to be present in a substantial share of entirely typical people. A gap is only informative when it is both statistically reliable, meaning larger than measurement error, and clinically unusual, meaning rare in the reference sample. Those are two separate tests and a difference has to pass both. Reports that discuss gaps without citing base rates are asking you to be impressed by something ordinary, and that omission is one of the clearest signals that a report was written to sound insightful rather than to be accurate.

What Changes When Subtests Move Online

Some subtest types transfer to computer administration cleanly, some transfer with modification, and some cannot transfer at all. Knowing which is which explains why online batteries look the way they do.

Transfers cleanly: matrix reasoning, figure weights, visual puzzles, symbol search, digit span with audio presentation. These are selection or short-response tasks where the computer measures more precisely than an examiner with a stopwatch.

Transfers with modification: block design becomes an on-screen assembly task, losing the observational information a clinician gets from watching hands, and gaining precise timing. Vocabulary and similarities become typed responses, which changes the task by adding a writing demand and removing the examiner's ability to probe an ambiguous answer.

Does not transfer: anything requiring clinical judgment during administration. Querying an incomplete response, deciding whether an answer reflects misunderstanding or genuine failure, and noticing that somebody has stopped trying are all part of standardised administration in a clinical battery and have no computer equivalent.

That last category is the real boundary between online and supervised assessment, and it is not about item quality. An online battery can use excellent items, time them precisely, and norm them properly, and still miss what an examiner notices in the first ten minutes. The honest framing is in Professional IQ Test vs Online IQ Test.

How ACIS Structures Its Subtests

ACIS administers twenty subtests across six domains, which is a deliberately wide structure for an online instrument and follows directly from the argument in section 1.

The six domains are verbal comprehension, fluid reasoning, visual spatial, working memory, processing speed, and quantitative reasoning. Each is measured by at least three subtests, so no index rests on a single task, and the quantitative domain is reported separately rather than folded into others for the reasons in section 8.

Every subtest states its time limit and discontinue rule before it begins, and the report presents index scores with their subtest composition visible, so that a reader can see what a domain score was built from rather than being handed a number with no provenance.

The reporting follows section 10. Index level differences are discussed. Subtest level differences are shown but not narrated into strengths and weaknesses, because the specificity is not there to support it. That is a deliberate restraint rather than an omission, and it is one of the clearest markers separating a psychometrically serious report from one written to be flattering.

Where ACIS is limited is section 11. There is no examiner, so no querying, no clinical observation, and no verification of conditions. Those limits are stated rather than glossed, and they are the reason an ACIS score is a well constructed estimate rather than a clinical assessment.

FAQ: Subtests and Test Structure

What is a subtest?

A block of items of one type, administered together and scored on its own scale. Several subtests combine into an index, and several indices into a full scale score.

Why not just use one really good test?

Because every task measures its target plus its own irrelevant demands. Only by combining tasks that share the target and differ in everything else does the target become isolable.

How many subtests does a proper battery use?

Typically ten to twenty. Fewer than about ten makes it hard to cover the broad abilities with more than one task each, which is the minimum for a defensible index.

Why do subtests use a mean of 10 instead of 100?

Convention. Subtests use mean 10 and standard deviation 3, indices use mean 100 and standard deviation 15. A subtest score of 13 equals an index of 115.

Which subtest is the best single measure of general ability?

Vocabulary and matrix reasoning typically rank highest, though which comes first varies by battery and sample. Neither is good enough alone to substitute for a full battery.

What is the difference between matrix reasoning and figure weights?

Both target fluid reasoning. Matrix items are inductive, requiring you to infer a rule. Figure weights items are deductive, with a single defensible answer, making them harder to solve by elimination.

Why is Arithmetic classified differently in different batteries?

Because it genuinely loads on several abilities. It demands working memory to hold the problem, quantitative knowledge to know the operations, and reasoning to select them.

What does digit span forward versus backward tell me?

Forward is closer to pure storage, backward adds manipulation. The gap is one of the few subtest-level comparisons with a defensible interpretation and published base rates.

Are supplemental subtests worse than core ones?

No. A subtest is supplemental for reasons of time, redundancy, or index composition. Figure Weights was supplemental in one battery generation and core in the next without changing.

What is subtest substitution?

Replacing a spoiled core subtest with a supplemental one from the same index. Manuals limit it strictly because norms were built on a specific combination.

Why should I not read into a single low subtest score?

Short measures are unreliable, subtest-specific reliable variance is small, and with sixteen subtests there are one hundred and twenty pairwise gaps, so a large one always exists.

What is subtest specificity?

The portion of a subtest's variance that is neither shared with general ability nor with its index. It is usually small and largely error, which is why single subtest interpretation is weakly supported.

Then what should I actually read in my report?

Index scores and the differences between them, with their confidence intervals. A twenty point gap between indices is worth discussing. A three point gap between subtests is not.

Why did batteries split perceptual reasoning into two indices?

Because reasoning about relationships and manipulating visual material are separable abilities that had been merged. Splitting them aligned the batteries with the CHC evidence base.

Are verbal subtests unfair to non-native speakers?

They measure ability as developed in a specific linguistic environment, so a second-language examinee scores below their capability. No statistical adjustment recovers it, which is why nonverbal indices are reported separately.

Why is working memory the domain most affected by test conditions?

Because anxiety and fatigue occupy the same limited capacity the task needs. A depressed working memory score in an anxious examinee is a finding about that session.

Do processing speed subtests measure motor skill?

Partly, and unavoidably. All of them require a physical response, so input device, fine motor control, and hand injury affect the score independently of cognitive speed.

Which subtests survive moving online?

Selection and short-response tasks transfer cleanly. Block design and open verbal responses transfer with modification. Anything needing clinical judgment during administration does not transfer.

Why does ACIS report quantitative ability separately?

Because Gq behaves like acquired knowledge rather than reasoning, and folding it into a fluid index hides a real quantitative deficit in somebody with strong reasoning and weak instruction.

How many subtests does ACIS use?

Twenty, across six domains, with at least three per domain so that no index rests on a single task.

Does a report need to say which subtests were administered?

Yes. A full scale score obtained with substitutions is weaker than the same figure from the standard combination, and the number itself does not reveal the difference.

Best Next Step

Subtests are the level at which testing is designed and the level at which it should not be interpreted. Understanding what each family targets tells you what a battery covers and where its blind spots are, and it tells you which numbers in your own report deserve attention.

If you want the domain structure that sits above the subtests, read Cognitive Domains. For the theoretical framework the whole hierarchy is built on, read The CHC Model. To see the structure in practice across twenty subtests, take the assessment.

Sources Behind This Page

Subtest composition and index structure are taken from publisher technical manuals. Claims about the weakness of subtest-level interpretation come from the peer reviewed factor analytic literature.

  • Pearson (2008). WAIS-IV Technical and Interpretive Manual. Subtest composition, core and supplemental designations, substitution rules and their limits.
  • Pearson (2014). WISC-V Technical and Interpretive Manual. Documents the separation of perceptual reasoning into distinct fluid reasoning and visual spatial indices, and the empirical criteria used for administration rules.
  • Pearson (2024). WAIS-5, Wechsler Adult Intelligence Scale, Fifth Edition. The current adult battery and its subtest and index structure.
  • McGrew, K.S. (2009). CHC theory and the human cognitive abilities project: standing on the shoulders of the giants of psychometric intelligence research. Intelligence, 37(1), 1-10. The broad and narrow ability taxonomy each subtest family is mapped onto.
  • Canivez, G.L., Watkins, M.W. & Dombrowski, S.C. (2016). Factor structure of the Wechsler Intelligence Scale for Children, Fifth Edition. Psychological Assessment. Finds the majority of reliable variance attributable to general ability rather than to the group factors, which is the technical basis for treating subtest-level interpretation cautiously.
  • Canivez, G.L. & Youngstrom, E.A. (2019). Challenges to Cattell-Horn-Carroll theory: empirical, clinical, and policy implications. Applied Measurement in Education. Reviews the evidence on how much interpretive weight index and subtest scores can carry.
  • American Educational Research Association, American Psychological Association & National Council on Measurement in Education. Standards for Educational and Psychological Testing. Requirements for reporting reliability and standard error at every score level a test reports, including subtests.
  • Voncken, L., Albers, C.J. & Timmerman, M.E. (2019). Improving confidence intervals for normed test scores. Behavior Research Methods. Open access. Why short subtests carry wide confidence intervals and composites carry narrower ones.
  • Crawford, J.R., Garthwaite, P.H. & Slick, D.J. (2009). On percentile norms in neuropsychology: proposed reporting standards. The Clinical Neuropsychologist, 23(7), 1173-1195. How subtest and index scores should be reported with uncertainty attached.
  • Pearson Clinical Assessment Scientific Council (2023). Standardized Clinical Assessment for Practitioners: A Primer. What standardised administration provides that unsupervised delivery cannot, including querying and clinical observation.
  • Buros Center for Testing. Independent test review body publishing the Mental Measurements Yearbook, which evaluates subtest composition and construct coverage in commercial batteries.
  • Riverside Insights. Woodcock-Johnson IV. The CHC-aligned battery that reports quantitative ability as a separate broad factor rather than folding it into other indices.
Take the assessment

You get a profile, not a number

ACIS measures six CHC domains across 20 subtests and reports each one with its own normed score and confidence interval, so you can see where you are strong and where you are not.

Free trial, no card required. Full report from $15.