Instrument Inventory

Types of IQ tests: the instruments, one row each

There are roughly a dozen intelligence batteries in serious professional use and a far larger population of online tests. This page lists them with the same fields for every row: publisher, edition, ages, administration mode, qualification level, what each is built for, and where each does not belong.

A boy with a pen looks up at a two color brain drawing labeled creative, idea and work.
Every figure in this inventory was read off a publisher page, a publisher flyer or a peer reviewed paper.

0 The Short Answer

The division that decides what an intelligence test score can be used for is not verbal against nonverbal, and not child against adult. It is supervised against unsupervised. Everything else on this page is a detail inside that split. A battery administered face to face by a qualified examiner, who verifies identity, controls the room, reads standardized instructions and watches effort, can support a decision about a person. A battery a person takes alone at a laptop cannot support that decision, no matter how many subtests it contains or how good its internal reliability looks, because the conditions under which the score was produced are unverified.

The honest complication is that this split does not map onto quality. Some unsupervised instruments are psychometrically careful and some supervised ones use norms collected more than twenty years ago. Supervision does not make a score accurate. It makes a score defensible, which is a different property, and it is the property that institutions require.

A second complication runs through the whole category. Several of the tests people call IQ tests are not intelligence batteries at all. The Wonderlic is a twelve minute hiring screen. The ASVAB is a military aptitude and classification battery whose owner publishes an explicit warning against using its scores for anything outside three validated purposes. Both get described as IQ tests constantly, and both are built for a different job.

This page is an inventory, not a recommendation. Every row carries the same fields, verified against the publisher's own page rather than repeated from a blog. The separate question of which instrument fits which decision belongs to its own page, and the question of what the individual tasks inside a single battery measure belongs to another. Neither is repeated here.

Level C

Pearson's top qualification tier, which gates the WAIS-5, the WISC-V and the WPPSI-IV.

Four modes

Administration modes defined by the International Test Commission, from open to managed.

Three uses

The only validated uses of an ASVAB score, according to the Office of People Analytics.

1 What Each Field In This Inventory Means

Six fields describe an instrument well enough to place it, and most pages that answer this query supply two of them. The fields used throughout this page are publisher, current edition and year, age range, administration mode, qualification level and intended purpose, plus a seventh that almost nobody publishes: where the instrument is not appropriate.

Publisher and edition year. Test publishing is concentrated. Pearson owns the Wechsler family, the Kaufman batteries, the Differential Ability Scales and the current Raven's. Riverside Insights owns the Woodcock-Johnson line and the Stanford-Binet. PAR owns the Reynolds scales. The edition year matters more than the name, because a normative update replaces the reference sample while keeping the items, and a new edition replaces both.

Age range. Published in years and months, because the boundaries are exact. A test standardized to 16 years 11 months has no norms at 17 years 0 months, and applying it there produces a number with no reference group behind it. Age bands are the whole mechanism by which raw performance becomes a standard score, so a range is a hard edge rather than a suggestion.

Administration mode. Individually administered by a qualified examiner, group administered under proctoring, or self administered without supervision. This is the field that decides what a score licenses.

Qualification level. Pearson's published qualifications policy defines three tiers. Level A products carry no special purchase requirement. Level B requires a master's degree in psychology, education, speech language pathology, occupational therapy, social work or counseling with formal training in ethical administration and interpretation, or an equivalent professional credential. Level C, in Pearson's own wording, requires a high level of expertise in test interpretation, and is open to holders of a doctorate in psychology or education with that formal training, or to those with state licensure or membership in bodies such as APA or NASP. Other publishers use their own schemes, and some publish no letter at all.

Purpose and non-purpose. The last two fields carry the weight. An instrument built to identify cognitive processing differences in a nine year old is not a worse test of adult ability, it is not a test of adult ability. Reading a row without those two columns is how people end up comparing an internal consistency coefficient across instruments that answer different questions.

Editions moveThree of the instruments below changed edition or norm set between 2018 and 2025. Any list of intelligence tests goes stale in about two years, which is why every row here names the publisher document it came from rather than asking you to trust the page.

2 Supervision Is the Line That Decides What a Score Can Do

The International Test Commission settled this vocabulary in 2005 and it remains the cleanest way to sort the whole category. Its International Guidelines on Computer-Based and Internet Delivered Testing, adopted by the ITC Council at its July 2005 meeting in Granada, define four modes of test administration. They are worth stating precisely, because the popular framing of this topic collapses all four into a single distinction between real tests and internet quizzes.

ITC modeHuman supervisionIdentity authenticatedTypical example
OpenNoneNoAn internet test with no registration at all
ControlledNoneNo, but access is restricted to known test takersAn online battery behind a purchase and a login
Supervised (proctored)Direct human supervision of conditionsYesAn examiner logs in the candidate and confirms proper administration
ManagedHigh supervision plus control of the environmentYesA dedicated testing centre with controlled access and equipment

Every instrument in this inventory sits in one of those four rows, and the row predicts the permissible use better than any reliability coefficient does. A supervised administration produces three things an unsupervised one cannot: a verified identity, a controlled and comparable environment, and an examiner's behavioural record of how the person actually worked. That third item is invisible in the score and does much of the interpretive work in a clinical report.

What unsupervised administration gives up is not precision in the narrow sense. An online battery can have excellent internal consistency and a clean factor structure. What it gives up is the chain of custody between the person and the number. Nobody can confirm the person tested was the person named, that no reference material was open in a second tab, that the room was quiet, or that the session was a first attempt rather than a fifth.

This is why the same score means different things depending on where it was produced, and why the question of whether an unsupervised score can be trusted has a two part answer: reliably measured, yes, potentially; institutionally usable, no. If you want a supervised administration, the practical routes are listed on the page about where supervised testing is actually available. If you want to understand how the unsupervised route works mechanically, that is covered under online administration.

3 The Wechsler Family: WAIS-5, WISC-V and WPPSI-IV

Three instruments cover the human lifespan between them, share an architecture, and are all gated at Pearson's Level C. They are the most widely used individually administered intelligence batteries in the English speaking world, and each is a separate standardization rather than a version of the others.

FieldWAIS-5WISC-VWPPSI-IV
Publisher and yearPearson, 2024Pearson, 2014Pearson, 2012
Ages16:0 to 90:116:0 to 16:112:6 to 7:7
Stated time45 minutes for the 7 subtest FSIQ; 60 minutes for the 10 primary index subtestsAbout 60 minutes for the core subtests30 to 45 minutes at ages 2:6 to 3:11; 45 to 60 minutes at ages 4:0 to 7:7
QualificationLevel CLevel CLevel C
ModeIndividually administered by a qualified examinerIndividually administered by a qualified examinerIndividually administered by a qualified examiner

All three figures above come from the publisher's own listings. Pearson's WAIS-5 product page gives the age band 16:0 to 90:11 and the two administration times, and lists five primary indexes plus the Full Scale IQ and a General Ability Index, alongside a longer set of ancillary indexes. The WISC-V listing gives ages 6:0 to 16:11 with the same five primary indexes. The WPPSI-IV listing gives ages 2:6 to 7:7 and, notably, produces the Fluid Reasoning Index only for the upper age band, because the tasks that carry Gf are not administrable at two and a half.

What these instruments are built for is individual diagnostic and eligibility work: intellectual disability determination, giftedness identification, the cognitive half of a learning disorder evaluation, and neuropsychological assessment after injury or illness. What they are not built for is screening at volume, self administration, or any use where a qualified examiner is not present. The five index structure maps onto the CHC model of cognitive abilities loosely rather than exactly, and the WAIS-5 in particular reorganized several index memberships relative to its predecessor. If you need that specific comparison, it lives on the page covering what changed between the WAIS-IV and the WAIS-5, and the instrument itself is described in more depth on the WAIS-5 reference page.

4 Stanford-Binet 5 and the KABC-II Normative Update

These two share nothing except individual administration and a place in the standard clinical toolkit, and they illustrate opposite design philosophies. The Stanford-Binet spans almost the entire lifespan in one instrument. The Kaufman battery deliberately restricts itself to childhood and offers two competing theoretical lenses on the same data.

The Stanford-Binet Intelligence Scales, Fifth Edition, authored by Gale Roid and published in 2003, is distributed by Riverside Insights and by Stoelting. The Stoelting product page gives an age range of 2 to 85 and above, ten subtests organized across three easel books, and roughly five minutes per subtest. It produces a Full Scale IQ, a Verbal IQ, a Nonverbal IQ, five factor indexes (Fluid Reasoning, Knowledge, Quantitative, Visual-Spatial and Working Memory) and an Abbreviated Battery IQ drawn from the two routing subtests. Its distinguishing property is range: the item set extends far enough in both directions to be usable with profoundly impaired examinees and with the far upper tail, which is why it still appears in gifted identification protocols decades after publication.

The Kaufman Assessment Battery for Children, Second Edition Normative Update, is the current Kaufman edition. Pearson's KABC-II NU flyer gives a publication date of 2018, an age range of 3 to 18, Qualification Level C, paper and pencil administration, and completion times of 25 to 60 minutes for the core battery under the Luria model or 30 to 75 minutes under the CHC model. The dual model is the point of the instrument. The examiner chooses, before testing, whether to interpret results through a Luria neuropsychological frame that excludes acquired knowledge, or through a CHC frame that includes it. The same child yields two different global composites depending on that choice, which is a design decision rather than an inconsistency.

Neither instrument is appropriate outside its band. The KABC-II NU has no adult norms at all. The Stanford-Binet has adult norms but a reference sample collected in the early 2000s, which is a real limitation on any absolute score claim. Structural detail on the Binet lives on the page describing how the Stanford-Binet 5 is built, and the direct instrument comparison is covered under the WAIS-5 against the Stanford-Binet 5.

Norm age is a field tooThe SB5 reference sample dates from 2003. That is not a defect in the instrument, but it does mean an SB5 standard score and a WAIS-5 standard score are anchored to populations tested two decades apart, and the two numbers are not interchangeable at the decimal level.

5 The CHC Batteries: Woodcock-Johnson and the DAS-II

Both of these instruments changed under the field's feet recently, and most published lists of IQ tests have not caught up. The Woodcock-Johnson is the battery most directly built on Cattell-Horn-Carroll theory, and the Differential Ability Scales is the one designed around out of level testing.

The Woodcock-Johnson IV, published by Riverside Insights in 2014, covers ages 2 to 90 and above across three co-normed batteries: Tests of Cognitive Abilities, Tests of Achievement and Tests of Oral Language. Co-norming is the feature that made it the standard choice for ability and achievement discrepancy work, because cognitive and academic scores come from the same reference sample rather than two unrelated ones. Riverside's own WJ IV page now states that WJ IV kits are available through 31 December 2026 and record forms through 31 December 2027, and describes the WJ IV normative data as more than a decade old.

The replacement arrived in February 2025. Riverside's Woodcock-Johnson V page describes an assessment system normed between 2022 and 2024 containing 20 cognitive tests across 12 clusters and 26 achievement tests across 22 clusters, plus a nonverbal battery, delivered on a fully digital platform with no physical kit shipped. Examiners need a laptop and an examinee tablet. That is a structural change, not a cosmetic one: it moves the instrument's administration requirements away from a suitcase of easels and towards a device deployment, and it is the single largest recent shift in this category.

The Differential Ability Scales-II received a narrower change. Pearson's DAS-II NU School-Age flyer gives a publication date of 2023, Qualification Level C, paper and pencil administration, an Early Years range of 2:6 to 6:11 and a School-Age range of 7:0 to 17:11, a core battery of 45 to 60 minutes plus 30 minutes of diagnostic subtests, and a school age normative sample of more than 500 cases. Pearson states plainly that the items, subtests and test structure did not change, so existing users need no retraining. That is the textbook definition of a normative update: same instrument, new reference population.

Both batteries are individually administered and both report profiles across multiple broad ability domains rather than a single number. That is the property that makes them useful when a cognitive profile is uneven, and useless as a quick screen.

6 Brief and Nonverbal Instruments: RIAS-2 NU and Raven's 2

Two instruments exist to buy back time or to strip out language, and both trade coverage for that. Neither is a shortened version of a longer battery. Both were designed from the start around a narrower target.

The Reynolds Intellectual Assessment Scales, Second Edition, Normative Update, is published by PAR. Its RIAS-2 NU product page gives an age range of 3 to 94, Qualification Level C, and administration times of 20 to 25 minutes for the intelligence assessment, 8 to 10 minutes for the memory assessment and 8 to 10 minutes for speeded processing. Its normative update is anchored to 2023 United States census data. It produces a Verbal Intelligence Index, a Nonverbal Intelligence Index, a Composite Intelligence Index representing general ability, plus a Composite Memory Index and a Speeded Processing Index. PAR's stated design goal is minimal motor demand, which makes it usable with examinees whose hand function would depress performance on a battery full of block manipulation. The cost of that design is coverage: a twenty five minute intelligence assessment cannot sample six broad domains the way a two hour battery can.

Raven's 2 Progressive Matrices, Clinical Edition, is Pearson's current Raven's. The Raven's 2 flyer gives a publication date of 2018, Qualification Level B rather than C, either paper and pencil or digital administration through Q-global, and administration individually or in small groups. Pearson's US listing gives an age span of 4 to 90, while the flyer describes examinees from four to eighty five, so the row below carries the listing figure and this sentence carries the discrepancy. Pearson states no completion time for Raven's 2 on its US, Canadian, UK or Asian store pages, an absence recorded alongside the 20 to 45 minutes the same publisher does state for the older Standard Progressive Matrices on the page collecting stated administration times. The instrument reports a raw score and a percentile rank against general population norms. It does not report a Wechsler style index profile, and the flyer notes that its harder items are concentrated in the top twenty percent of the population, which is a deliberate ceiling decision.

Raven's 2 is the instrument people mean when they say a test is language free, and the claim needs qualifying. Reduced language load is real and useful with English language learners and with communication disorders. It does not make a matrices score equivalent to a broad ability composite, because a single task family samples one slice of ability however reliably it does so. That distinction sits at the centre of both nonverbal ability testing and culture reduced testing, and the instrument itself is described in detail on the Raven's 2 profile page.

Brief is not the same as weakA short instrument can be highly reliable for the construct it targets. What it cannot be is broad. A composite built from two indexes and a composite built from twenty subtests are different measurements even when both are labelled general ability.

7 The Full Inventory, One Row Per Instrument

This is the whole list, with identical fields for every row including the unsupervised ones. Where a publisher does not state a qualification level, the cell says so rather than guessing. Where an edition has been superseded, both editions appear, because practitioners are still administering the older one.

InstrumentPublisher, edition yearAgesAdministration modeQualificationBuilt forNot appropriate for
WAIS-5Pearson, 202416:0 to 90:11Individual, qualified examinerLevel CAdult FSIQ and five primary indexes for clinical and eligibility workGroup screening, self administration, volume testing
WISC-VPearson, 20146:0 to 16:11Individual, qualified examinerLevel CSchool age cognitive profile and special education eligibilityAdults, preschoolers, any unsupervised delivery
WPPSI-IVPearson, 20122:6 to 7:7Individual, qualified examinerLevel CPreschool and early primary cognitive evaluationAnyone past 7:7; no Gf index below age 4
Stanford-Binet 5Roid, 2003; Riverside Insights and Stoelting2 to 85 and aboveIndividual, qualified examinerRestricted purchase; distributor page states no letter levelVery wide span, including the extreme upper and lower rangesAbsolute comparison against instruments normed two decades later
KABC-II NUPearson, 20183 to 18Individual, paper and pencilLevel CChild cognitive processing under either a CHC or a Luria interpretationAdults; cross model comparison of the two global composites
DAS-II NU School-AgePearson, 2023Early Years 2:6 to 6:11; School-Age 7:0 to 17:11Individual, paper and pencilLevel CChild ability profile with out of level testing across batteriesAdults; group settings
Woodcock-Johnson IVRiverside Insights, 20142 to 90 and aboveIndividual, qualified examinerRestricted purchaseCo-normed cognitive, achievement and oral language batteriesBeing treated as current; kits end 31 December 2026
Woodcock-Johnson VRiverside Insights, February 20254 to 90 and above; normed 2022 to 2024Individual, fully digital, examiner administeredRestricted purchaseCurrent WJ edition: 20 cognitive tests across 12 clustersPaper only settings; unsupervised delivery
RIAS-2 NUPAR, normative update on 2023 census data3 to 94Individual, qualified examinerLevel CBrief general ability with verbal, nonverbal, memory and speed indexesA full six domain profile; broad ability inference from 25 minutes
Raven's 2 Clinical EditionPearson, 20184 to 90 on the US listingIndividual or small group; paper or Q-globalLevel BNonverbal reasoning estimate with reduced language demandStanding in for a broad composite; it reports raw score and percentile
Wonderlic cognitive testWonderlic, current Select platformWorking age applicantsOnline, employer deployed, typically unproctoredEmployer purchase; no clinical level publishedTwelve minute pre-hire cognitive screen, 50 itemsAny interpretation as an intelligence score
ASVAB and the AFQTUS Department of DefenseMilitary applicants and secondary studentsGroup; CAT-ASVAB proctored, PiCAT unproctoredGovernment administeredEnlistment selection, job classification, career explorationEvery other purpose, per the Office of People Analytics
ICARCondon and Revelle, 2014; public domainAdults in research samplesOnline, unproctored, self administeredNone; items are public domainResearch needing a cognitive ability variable at scale and no costIndividual score reporting; there is no clinical norm table
ACISacisiq.com, technical manual version 1.4, August 202616 to 90Online, unsupervised, self administeredNonePersonal cognitive profile: 20 subtests, FSIQ and six primary indexesDiagnosis, accommodations, hiring, court, society admission
Open web IQ quizzesVarious; usually no publisher of recordUnstatedOpen mode: no registration, no supervisionNoneEntertainment and lead captureAny score claim at all; typically no manual and no published norms

Read down the administration column before reading anything else. Ten of these fifteen rows require a person in the room. Four do not, and one is delivered both ways depending on which ASVAB variant a candidate sits. That column, not the publisher column and not the age column, is what determines whether a resulting number can appear in a school file, a clinical report or a court exhibit.

8 Aptitude Screens Are Not Intelligence Batteries

The Wonderlic and the ASVAB are described as IQ tests more often than any actual intelligence battery, and neither publisher makes that claim. The distinction is not pedantry. An intelligence battery is built to estimate broad cognitive ability and to report a profile against age norms. An aptitude or trainability screen is built to predict performance in a specific setting, and it is validated against that outcome rather than against a general ability construct.

Wonderlic's own cognitive ability page describes a 50 question test with a 12 minute limit, covering simple math, basic logic, language comprehension, spatial reasoning and pattern identification, with items that grow harder through the test. The framing throughout is prediction of job performance: identifying best fit candidates, deciding who receives an offer, and tailoring onboarding. Nowhere does the publisher present the result as an intelligence quotient, and the score is weighted against personality and motivation measures inside the wider Select platform rather than reported as a standalone ability estimate.

The ASVAB case is stronger still, because its owner published a document specifically to stop people misusing it. The official subtest listing describes ten tests across four domains: Word Knowledge and Paragraph Comprehension for verbal, Arithmetic Reasoning and Mathematics Knowledge for math, five tests for science and technical content, and Assembling Objects for spatial. An executive note from the Office of People Analytics titled Appropriate Use of ASVAB Scores states that only three uses are validated: selection into the military using the AFQT, classification into occupations using service specific composites, and career exploration. The same document lists the misuses it warns against, including evaluating educational outcomes, predicting SAT or ACT performance, and making postsecondary admission or scholarship decisions.

The AFQT itself is a composite of four subtests (Arithmetic Reasoning, Mathematics Knowledge, Paragraph Comprehension and Word Knowledge), reported as a percentile from 1 to 99 against a sample of 18 to 23 year olds tested in a 1997 national norming study. A percentile ceiling of 99 cannot resolve anything above roughly the top one percent, and the composite reports no domain profile. The arithmetic of translating an AFQT into IQ units, and what that translation does and does not license, is worked through on the page comparing the AFQT against IQ. The broader category distinction is covered under aptitude against intelligence, and workplace instruments generally under hiring assessments.

9 Unsupervised Instruments: Public Domain, Research and the Open Web

Below the supervised line the population splits three ways, and lumping them together is the most common error in this category. There are public domain research instruments, commercial online batteries, and open web quizzes with no publisher of record. They differ enormously in evidence and not at all in administration mode.

The International Cognitive Ability Resource is the reference case for the first group. David Condon and William Revelle described it in Intelligence, volume 43, pages 52 to 64, in 2014, as a public domain measure built for large scale remote data collection. The initial validation covered four item types: 9 Letter and Number Series items, 11 Matrix Reasoning items, 16 Verbal Reasoning items and 24 Three-dimensional Rotation items. The paper reports alpha of .93 and omega total of .94 for the full 60 item set, with alpha of .81 and omega total of .83 for the 16 item sample test, and reports administration to roughly 97,000 online participants. The authors argue that public domain status did not compromise validity at that scale, while acknowledging that unproctored administration offers no safeguard against the use of external resources.

That combination is exactly what makes ICAR useful and exactly what limits it. It is free, it is modifiable, it can be dropped into any survey, and it produces a defensible cognitive ability variable for research. It does not produce an individual score a person can interpret about themselves, because there is no norm table of the kind a commercial publisher maintains and no interpretive report. The founding sample shows why: 96,958 self selected internet volunteers, 66 percent of them female, median age 22, and 78.1 percent located in the United States, a composition that supports item level analysis rather than a percentile for one person, as the dedicated account of the ICAR 16 and ICAR 60 forms sets out.

The third group is the largest by traffic and the emptiest by evidence. Open mode quizzes require no registration, publish no manual, name no normative sample, and frequently report a number scaled to flatter. Their existence is why the entire unsupervised category carries a reputational discount, including for instruments that document themselves properly. The distinction between the two is the subject of free quizzes measured against validated instruments, with the practical landscape covered under the serious online batteries.

Two adjacent categories are worth naming so they are not mistaken for this one. Admission tests used by high IQ societies are usually supervised or use supervised evidence, and high range tests are untimed unsupervised instruments built for the far tail with normative samples of self selected volunteers. Neither is a general purpose ability battery, and Mensa admission testing in particular is a pass or fail threshold rather than a profile.

Public domain has a costICAR items are published, searchable and reusable by design. That is a feature for researchers who need to know exactly what was administered, and a permanent security limitation for anyone who wanted to use the same items in a setting where the score mattered.

10 Where ACIS Sits, With the Same Fields

ACIS is an unsupervised online adult battery, and it belongs in this inventory on the same terms as everything else. Publisher: acisiq.com. Current documentation: technical manual version 1.4, updated 3 August 2026. Ages 16 to 90. Administration mode: self administered online, with no examiner present and no identity authentication, which places it in the International Test Commission's controlled mode rather than open mode, because access is tied to a purchase and an account. Qualification level: none, because there is no gate to purchase.

Structure: 20 subtests across six primary cognitive domains, producing a Full Scale IQ plus six primary indexes, with subtest scaled scores at mean 10 and standard deviation 3 and composites at mean 100 and standard deviation 15. Three forms exist. Quick administers 6 subtests, Optimized administers 13, and Full Scale administers all 20. The forms are not interchangeable: a Quick administration does not cover every domain and cannot be read as a Full Scale result.

Reported evidence, from the published manual rather than estimated here: composite omega of .9886 for Full Scale IQ, a g loading of .958 for that composite, and a standard error of measurement of about 1.60 IQ points. The higher order g confirmatory model reports CFI .9761, TLI .9726, RMSEA .0406, SRMR .0217, and a chi-square of 916.703 on 166 degrees of freedom. The adult reference frame contains 3,243 English speaking records across ages 16 to 90, and the technical analysis set behind the factor and reliability tables is 2,750 complete records.

The limitations belong in the same paragraph as the figures. That reference frame is self selected rather than census sampled: people who choose to buy a cognitive assessment are not a random draw from the adult population, and the frame is a modelled adult reference rather than a stratified national sample. Administration is unsupervised, so nothing verifies who sat the test or under what conditions. And ACIS is not a clinical instrument. It is not appropriate for diagnosis, for disability or accommodation documentation, for hiring decisions, for legal or forensic use, or for high IQ society admission. Those uses require the supervised column of the table above.

A separate business line, ACIS Professional, exists for research, screening and educational contexts, with participant links that collect no personal data and verified score export. It is described at the research workspace landing page and pitched on the home page section for people who administer the assessment to participants. The score architecture, including the fifteen point standard deviation the composites use and what the detailed report contains, is documented rather than summarized here.

11 What This Inventory Leaves Out, and Why

Several whole categories were excluded on purpose, and naming them is more useful than silently omitting them. Each is a real family of instruments that people find while searching for types of IQ tests, and each fails at least one condition for belonging on the list above.

Group administered school ability tests. Districts screen entire grade levels with paper or online ability tests delivered to a room at once. They are proctored, so they are not unsupervised, but they are not individually administered either, and their purpose is placement screening rather than individual diagnosis. A group ability score is a reasonable trigger for a full evaluation and a poor substitute for one. Where that screening leads is covered under the gifted threshold and, for grown examinees, adult giftedness assessment.

Abbreviated forms of full batteries. Publishers sell two and four subtest short forms of their own instruments. These are not separate tests, they are documented reductions of an existing standardization, and their manuals state the wider confidence intervals that follow from fewer items. Listing them as instruments would double count the parent battery.

Cognitive screening tools used in medicine. Brief bedside instruments that detect probable cognitive impairment are not intelligence tests and do not produce IQ scores. They are pass or refer screens with clinical cut points, calibrated to identify decline rather than to locate a person on an ability distribution. Confusing the two is a category error with real consequences, because a screening cut point says nothing about premorbid ability.

Academic admission exams. The SAT, ACT, GRE, GMAT and LSAT correlate with cognitive ability without being measures of it. They are curriculum sensitive, coachable by design, and normed on applicant pools rather than on the general population.

Single subtest tasks sold as tests. A matrices set, a digit span task or a symbol substitution task is one measurement, not a battery. This is where the difference between an instrument and a task matters, and it is why a comprehensive battery and a single reasoning task produce numbers that should not be compared. The reasons a battery uses many tasks instead of one, including how per item time limits shape what each task captures, are handled on their own pages.

12 How to Check Any Row Yourself, and Why Names Drift

Test names are sticky and test contents are not, which is why so many published lists carry figures that were correct a decade ago. Six of the instruments above changed edition or norm set since 2018. Anyone quoting a figure about an intelligence test should be able to say which edition it applies to, and the honest way to get there is short.

Start with the distinction that causes the most confusion. A normative update keeps the items, the subtests and the structure, and replaces only the reference sample. Pearson said this explicitly about the DAS-II NU: existing customers need no additional training. A new edition changes the item content, often the subtest list and sometimes the index structure. The Woodcock-Johnson V is not a digital conversion of the Woodcock-Johnson IV, and the WAIS-5 is not a reprint of the WAIS-IV. Scores across a normative update are comparable in kind and shifted in level. Scores across an edition boundary are a different measurement.

Then check three things on the publisher's own page rather than on any secondary source, this one included. First, the publication date, which is usually stated on the product listing or in a downloadable flyer. Second, the qualification level, which tells you the administration model the publisher assumes. Third, the age range in years and months, which tells you whether norms exist for the person in front of you. If a claim about a test cannot be traced to one of those three, it is folklore.

Norm age is the field people skip and the one that moves scores most. Reference samples drift relative to the population, which is why publishers reissue norms without touching items at all, and why the secular trend in tested performance is a practical publishing problem rather than an academic curiosity. An instrument normed in 2003 and an instrument normed in 2024 will not agree on the same person, and the disagreement is not evidence that either one is broken.

Finally, note what a test name does not tell you. It does not tell you the construct coverage, the reliability of the composite you care about, or whether the score you are being shown is a standard score, a percentile or a raw total. Raven's 2 reports a raw score and a percentile rank. The Wechsler batteries report standard scores. Those are different quantities with different properties, which is the whole subject of the difference between a score and a percentile, and it feeds directly into what accuracy means for a test. Nor does a name fix the words printed beside the number, since Pearson's WAIS-IV sample report labelled its top band Very Superior and the 2024 WAIS-5 report replaced that vocabulary with directional descriptors, while the only manual that ever carried a genius row was Terman's of 1916, a sequence traced on the page following where the genius label entered and left the manuals.

One useful habitWhen you read any figure about an intelligence test, ask which edition it belongs to and what year its norms were collected. Those two questions eliminate most of the incorrect claims circulating about this category, including several that appear on pages ranking above this one.

13 What the Testing Standards Say About the Whole Inventory

The organising principle of this page is not an opinion about supervision. It is the position taken by the field's governing document. The Standards for Educational and Psychological Testing, published jointly in 2014 by the American Educational Research Association, the American Psychological Association and the National Council on Measurement in Education, define validity as the degree to which evidence and theory support the interpretations of test scores for proposed uses. That definition does something specific: it moves validity out of the instrument and into the claim.

The consequence for an inventory like this one is direct. There is no such thing as a valid intelligence test, only evidence supporting particular interpretations of particular scores for particular purposes. The Office of People Analytics applied exactly this reasoning to the ASVAB, stating that a test can be neither valid nor invalid because validity is not a property of a test, and that each new proposed use must undergo its own validation. That is why the final column of the table above, the one naming where an instrument does not belong, is psychometrically the most important column and the one almost no competing list publishes.

It also explains why administration mode carries so much weight. Interpretation for a high stakes use requires evidence collected under conditions resembling the use. A score produced without identity verification, without a controlled environment and without a behavioural record supports a narrower set of interpretations than a score produced with all three, regardless of the internal consistency either one reports. The International Test Commission built its four mode framework around precisely that observation, and APA's testing standards resources point practitioners at the same principle.

Applied to this site's own instrument, the boundary is stated rather than implied. ACIS is unsupervised, its reference frame is self selected rather than census sampled, and it is not a clinical or diagnostic instrument. It supports interpretations about a person's own cognitive profile for that person's own use. It does not support diagnosis, accommodation documentation, employment decisions, forensic use or society admission, and no reliability coefficient would change that, because the limiting factor is the administration mode and not the measurement.

Read the inventory that way and the list stops being a menu. It becomes a map of which claims each instrument can carry, which is the only question about types of IQ tests that has a stable answer. If you want the interpretive layer on top of a score, that lives on the pages covering what a score means in practice and how rare a given score is.

14 Frequently Asked Questions

What are the main types of IQ tests?

They divide first by administration mode: individually administered batteries given by a qualified examiner, group administered tests given under proctoring, and self administered online tests with no supervision. Within each group they divide again by age band and by whether they report a broad profile or a single narrow score.

Are all IQ tests the same?

No. They differ in normative sample, age coverage, construct breadth, administration mode and purchase restriction. Two instruments can both report a mean of 100 and a standard deviation of 15 and still be anchored to populations tested twenty years apart, which makes their numbers non interchangeable.

Which IQ test do psychologists actually use with adults?

In English speaking practice, most often the WAIS-5, published by Pearson in 2024 for ages 16:0 to 90:11 at Qualification Level C. The Stanford-Binet 5, the Woodcock-Johnson and the RIAS-2 NU also carry adult norms and appear in specific referral contexts.

What does a qualification level mean when buying a test?

It is the publisher's purchase restriction. Pearson defines Level A as unrestricted, Level B as requiring a relevant master's degree or professional credential with training in ethical administration, and Level C as requiring a doctorate or licensure plus high expertise in interpretation.

Is the Wonderlic an IQ test?

Wonderlic does not present it as one. The publisher describes a 50 question, 12 minute cognitive test framed as a predictor of job performance, weighted alongside personality and motivation measures inside its hiring platform rather than reported as a general ability score.

Is the ASVAB an IQ test?

No. It is a military aptitude battery. Its owner, the Office of People Analytics, publishes a note stating that only three uses are validated: enlistment selection, occupational classification and career exploration, and it explicitly warns against using ASVAB scores for other purposes.

What is the difference between the AFQT and the ASVAB?

The ASVAB is the full battery of ten subtests across verbal, math, science and technical, and spatial domains. The AFQT is a composite of four of them, Arithmetic Reasoning, Mathematics Knowledge, Paragraph Comprehension and Word Knowledge, reported as a percentile from 1 to 99.

Which IQ test covers the widest age range?

The Stanford-Binet 5 runs from age 2 to 85 and above, and the RIAS-2 NU runs from 3 to 94. The Woodcock-Johnson IV covers 2 to 90 and above. No single instrument covers every age with equally recent norms.

What is a normative update and how is it different from a new edition?

A normative update replaces the reference sample while keeping the items, subtests and structure unchanged, so no retraining is needed. A new edition changes item content and often the subtest list and index structure, which makes scores across the boundary a different measurement.

Is Raven's 2 a full IQ test?

It is a nonverbal reasoning instrument. Pearson's flyer describes paper or digital administration at Qualification Level B, individually or in small groups, reporting a raw score and a percentile rank rather than a Wechsler style index profile across multiple cognitive domains.

What is the ICAR?

The International Cognitive Ability Resource is a public domain item pool described by Condon and Revelle in Intelligence in 2014. It contains letter and number series, matrix reasoning, verbal reasoning and three dimensional rotation items, and was built for large scale unproctored research rather than individual reporting.

Why does supervision matter more than the number of subtests?

Supervision supplies identity verification, controlled conditions and a behavioural record of effort. Without those, a score can be internally reliable and still unusable for any institutional decision, because nothing connects the number to a specific person tested under known conditions.

Can an online IQ test be used for a diagnosis?

No. Diagnostic use requires an individually administered battery given by a qualified examiner under standardized conditions, plus history, observation and convergent evidence. An unsupervised score, however carefully constructed, cannot supply the administration conditions that a diagnostic interpretation depends on.

Which tests are used to identify giftedness in children?

Individually administered batteries with adequate ceilings, most often the WISC-V, the Stanford-Binet 5 or the WPPSI-IV at preschool age. Group ability screens are commonly used first to decide who receives a full individual evaluation, not to make the identification itself.

What changed with the Woodcock-Johnson in 2025?

Riverside Insights released the Woodcock-Johnson V in February 2025, normed between 2022 and 2024, with 20 cognitive tests across 12 clusters and fully digital administration requiring an examiner laptop and an examinee tablet. WJ IV kits remain available through the end of 2026.

Do any intelligence tests avoid language entirely?

None avoid it entirely, but some reduce it substantially. Raven's 2 minimizes language demand through matrix items, and the KABC-II NU offers a Luria interpretation that excludes acquired verbal knowledge from the global composite. Reduced language load narrows construct coverage as well.

What is the difference between an intelligence battery and an aptitude test?

An intelligence battery estimates broad cognitive ability and reports a profile against age norms. An aptitude screen predicts performance in a defined setting and is validated against that outcome. The same person can score differently on each because the two answer different questions.

How current does a test's normative sample need to be?

There is no fixed rule, but publishers reissue norms precisely because reference samples drift relative to the population. Riverside describes the Woodcock-Johnson IV norms as more than a decade old, which is the kind of statement worth checking before quoting any absolute score.

Where does ACIS fit in this list?

As an unsupervised online adult battery for ages 16 to 90, with 20 subtests across six domains, no purchase qualification, and a self selected reference frame. It supports personal profile interpretation only, not diagnosis, accommodations, hiring, forensic use or high IQ society admission.

Why do published lists of IQ tests disagree with each other?

Because most repeat figures from other lists rather than from publisher pages, and because editions change faster than articles are updated. Six instruments in this inventory changed edition or norm set since 2018, and stale entries survive indefinitely once copied.

Which single field should I check first when comparing two tests?

Administration mode. It determines what claims a resulting score can carry, and it is more decisive than publisher, subtest count or reported reliability. Age range and norm year come next, because together they decide whether the score has a valid reference group at all.

Take the assessment

You get a profile, not a number

ACIS measures six CHC domains across 20 subtests and reports each one with its own normed score and confidence interval, so you can see where you are strong and where you are not.

Free trial, no card required. Full report from $15.