The adaptive IQ test, explained: how a test picks its questions, when it stops and what it costs
An adaptive IQ test chooses each question from the answers you have already given, so a short test can measure precisely near your level. This page explains how that works, from item response theory to stopping rules and item exposure, works the arithmetic on what it saves and what it costs, and sets out what the GRE, CAT-ASVAB, NIH Toolbox, PROMIS and several online tests publish about their design.
The illustration is decorative; the CAT-ASVAB, a real adaptive battery, lists 15 scored questions for most of its ten subtests on the official ASVAB site.
0 The short answer
An adaptive IQ test is a test that picks the next question from your earlier answers, using a statistical model of item difficulty, and stops when a rule says the estimate is precise enough or a fixed number of items is reached. The saving is real but narrower than the label suggests: it is largest at the very high and very low ends of the scale, it depends on a large calibrated item bank and on exposure control, and it changes nothing about whether the score is normed or valid. The word "adaptive" covers very different designs, from the GRE, which adapts once per measure by section, to tests that stop on a precision rule, so what matters is the published selection rule, stopping rule and bank, not the label.
DisclosureACIS sells a competing paid assessment, so we have a commercial interest in how this page reads. ACIS is not an adaptive test, and this page says so where it matters. Every product fact below comes from the publisher or standards document named in the text, read on October 6, 2026 unless another date is given, and every figure we derived ourselves is labeled as our arithmetic.
15
Scored questions listed for most CAT-ASVAB subtests on the official ASVAB site; the technical bulletin says each test ends after a fixed number of items or the time limit.
0.3
The standard error at which the NIH Toolbox picture vocabulary and reading recognition tests stop, up to a maximum of 25 items, per Gershon and colleagues in 2014.
4 to 12
Items administered by a PROMIS computerized adaptive test, per HealthMeasures.
1 What Is an Adaptive IQ Test, and What Exactly Does It Adapt?
An adaptive test changes what it asks, how long it runs, or both, according to how the test taker has performed so far, and the three main designs adapt at very different points. The International Test Commission and the Association of Test Publishers define adaptive testing in their 2022 Guidelines for Technology-Based Assessment as the personalized delivery of assessments with optimized precision in estimating ability. They name two ways to personalize: adapt the number of items, so that some people take shorter or longer tests, or adapt the nature of the items, typically by matching difficulty to ability. They also separate item-level designs, which adapt after every item, from multistage designs, which adapt after pre-designated blocks of two or more items.
That distinction explains most of the confusion in product listings. The GRE General Test is "section-level adaptive" in the words of ETS, whose page on the test structure we read on October 6, 2026: the first section of each measure, Verbal and Quantitative, is of average difficulty, and the difficulty of the second section depends on performance on the first. In the ITC vocabulary that is a multistage design with two stages, which is our classification, not ETS's wording. The CAT-ASVAB, the computerized form of the military's aptitude battery, adapts item by item. A test that moves up one difficulty level after each correct answer and down one after each error is also adaptive, but by a simpler rule that does not use a statistical model to choose items. All three deserve the word, and they behave differently, which is why a bare "this test is adaptive" tells a buyer very little.
An adaptive test has parts. The ITC guidelines list them, citing Kingsbury and Weiss (1984) and Luecht (2016): a large and fully calibrated item bank where every item has complete and accurate metadata, a starting point or initial ability estimate, an item selection algorithm, a scoring algorithm and one or more termination criteria. Content or exposure constraints are often added to the selection step. The same document says these components and their parameters should be documented. The Standards for Educational and Psychological Testing (AERA, APA and NCME, 2014), the profession's joint reference on test quality, make the same demand in Standard 4.3, quoted in the section on checking an adaptive claim.
Adaptive and short are separate properties. A test can be short and fixed, as the page on what a short battery costs you explains, and a test can be adaptive and long.
2 How Can Two People Who Saw Different Questions Get Comparable Scores?
Item response theory (IRT) makes adaptive scores comparable by placing every item and every person on one common scale, so the score depends on which items were answered correctly and how hard they were, not on how many. In a fixed test, a raw score such as 31 of 40 means the same thing for everyone because everyone saw the same 40 items. In an adaptive test two people may share almost no items, so counting correct answers would be meaningless: half right on hard items is more than half right on easy ones. The Standards put the resolution plainly: two examinees may rarely if ever receive the same set of items, yet adaptive test scores can be reported on a common scale and function much like scores from a single alternate form of a test that is not adaptive.
The procedure that does the placing is IRT, covered in full in textbook treatments such as Lord's Applications of Item Response Theory to Practical Testing Problems (1980, reissued 2012), the Wainer and colleagues primer, Computerized Adaptive Testing (2000), and van der Linden and Glas's edited volume, Elements of Adaptive Testing (2010). Every specific figure on this page comes from the primary documents listed under Sources; the textbooks supply the standard background. Each item gets parameters estimated from a large calibration sample, and each person gets an ability estimate, usually written theta, on the same scale. The Rasch or one-parameter model gives items a difficulty. The two-parameter logistic (2PL) model adds a discrimination parameter, which says how sharply an item separates people just below its difficulty from people just above it. The three-parameter model adds a guessing floor for multiple choice items.
Real instruments differ in which model they declare. The ASVAB Technical Bulletin No. 1, published by the Defense Manpower Data Center in March 2006, says the three-parameter logistic model was selected for the CAT-ASVAB primarily for its mathematical tractability and its accuracy in modeling the response probabilities of multiple-choice questions. The NIH Toolbox picture vocabulary test, by contrast, was calibrated to the one-parameter Rasch model: Gershon and colleagues (2014) report scoring and calibrating the items with Winsteps software. BRGHT's method page says its twenty difficulty levels derive from Rasch and 2PL models fitted to the results of over 500,000 test takers.
Calibration is the expensive, invisible step. Before the NIH Toolbox picture vocabulary test could adapt, its developers tried out 625 candidate words across 21 forms of 40 to 60 items each, so that every word drew at least 200 responses, and the calibration sample was 4,703 paid participants from an online panel company, per Gershon and colleagues. Twenty-three words were removed for poor statistical behavior, leaving a 602-word bank. The Standards add that when pools are updated, statistical procedures must link new items to the existing IRT scale so that scores from alternate pools stay interchangeable. A bank that grows with unlinked items, or items calibrated on an unrepresentative sample, can drift without any visible change on screen.
3 How Does an Adaptive Test Choose the Next Question?
The standard rule is to give the item that will tell the test the most about the current ability estimate, which in practice means an item whose difficulty sits close to that estimate. The ASVAB Technical Bulletin states the logic directly: adaptive tests achieve maximum precision when each item administered is the most informative for the current estimate of the examinee's ability level. The information an item supplies is a mathematical function of the model. For a 2PL item, the standard textbook expression is the squared discrimination multiplied by the probability of a correct answer and the probability of an incorrect one, which peaks when those two probabilities are equal, that is, when the person has a 50 percent chance of success. Items far too easy or far too hard contribute almost nothing. The Standards put the same point in words: it has long been recognized that little is learned from examinees' responses to items that are much too easy or much too difficult.
The NIH Toolbox shows the logic at its simplest. Gershon and colleagues describe successive items following from a continually updated estimate of the respondent's ability, with test difficulty controlled so that the participant has a 50 percent likelihood of answering each item correctly. That is the maximum information rule for a Rasch model, in which all items discriminate equally. The official ASVAB page describes the same behavior in everyday terms: an incorrect answer leads to an easier item, and a correct answer leads to a harder one, selected from a pool that ranges from very easy to very hard.
Here is the arithmetic of why targeting matters. This is our arithmetic using the standard 2PL formula, not a figure from any of the sources, and it assumes every item has a discrimination of 1.5, a value we chose for illustration. Take a person whose ability is exactly average on the scale, theta of 0.
Item difficulty relative to the person
Information per item (our arithmetic)
Share of an on-target item
Equal (0 standard deviations away)
0.563
100 percent
1 standard deviation away
0.336
60 percent
2 standard deviations away
0.102
18 percent
3 standard deviations away
0.024
4 percent
An item two standard deviations off target tells the test less than a fifth as much as an item at the person's level, and one three standard deviations away tells it almost nothing. A fixed form must include easy and hard items so that it can serve everyone, so most of its items are off target for any single person. An adaptive test spends each answer where the information is.
Pure maximum information selection has a known flaw, and real systems add content and exposure constraints to contain it. The comment to Standard 4.3 in the Standards says that if an adaptive test is meant to measure several content subcategories, item selection procedures should ensure the subcategories are adequately represented, and the ITC guidelines list the same constraint family.
Scoring is the other half of selection. Because the item sequence differs, each score is computed from the response pattern against the item parameters. The CAT-ASVAB technical bulletin explains that its final estimate is the Bayesian mode, because that estimator is unaffected by the order of items and is slightly more precise than the sequential Owen's estimate. It also states that Bayesian estimators draw the estimate toward the mean of the prior, that this bias is inversely related to test length, and that it is larger for short adaptive tests. For a high or low scorer, a very short adaptive test therefore pulls the estimate toward the middle, which is one reason the stopping rule and minimum length matter.
4 When Does an Adaptive Test Stop, and Why Does That Rule Matter So Much?
The stopping rule decides how long the test runs and how precise the score is, and the two common choices, a fixed number of items or a target precision, trade equal length for equal precision. The Standards state both options in their chapter on test design: when tests are administered adaptively, test length is determined by stopping rules, which may be based on a fixed number of test questions or on a desired level of score precision. Neither is free. A fixed length gives everyone the same number of items, so precision varies with how well the bank covers a person's level. A precision rule gives everyone the same target, so length varies, and people whose level the bank covers poorly can be given very long tests.
The CAT-ASVAB uses a fixed length, and its designers say why. The ASVAB Technical Bulletin No. 1 states that each test ends after the examinee has completed a fixed number of items or reaches the time limit, whichever occurs first. Its rationale is that simulation studies have shown fixed-length testing to be more efficient than variable-length testing, because highly informative items are typically concentrated over a restricted range of ability, so examinees outside that range tend to receive long tests in which each added item contributes very little. On the official ASVAB site, most of the ten subtests list 15 scored questions and three list 10, which is that fixed length in practice.
The NIH Toolbox language tests use a precision rule with a cap. Gershon and colleagues write that most stopping rules are based on a pre-specified level of precision (variable length), a given number of items (fixed length) or a combination, and that the picture vocabulary test continues until the standard error of performance is below 0.3, with a maximum of 25 items, typically five minutes or less. They report that the 0.3 cutoff was chosen because more than 95 percent of subjects reached that accuracy in under five minutes, a limit set by the Toolbox design team. The oral reading recognition test stops at the same 0.3 or at 25 items, with a median of 20 items in 4 minutes. In the validation study the authors used a fixed 25-item version, mainly to collect more responses per item. A published item count is therefore worth reading with care: it may be the cap, the average or the validation design.
PROMIS combines the same ingredients. HealthMeasures, which maintains PROMIS, says on its page on computer adaptive tests that items are dynamically selected from an item bank based on a person's previous answers, that a CAT administers between 4 and 12 items, and that it stops under rules such as a minimum and maximum number of items. Luijten and colleagues (2025) give the default rule applied in the data they analyzed from the Dutch-Flemish PROMIS assessment center: a minimum of 4 and a maximum of 12 items, or a standard error of the ability estimate below 0.32.
In the IRT framework, Luijten and colleagues explain, reliability can be expressed as the standard error of theta, which is directly related to total test information, and they give 0.32 as a commonly recommended stopping value for assessing individuals, corresponding to a reliability of 0.90. Our arithmetic, assuming an ability scale with a variance of 1, agrees: one minus 0.32 squared is 0.898, and one minus 0.30 squared is 0.91. If one unit of the ability scale is 15 IQ points, a standard error of 0.32 is 4.8 points, and a 95 percent interval of plus or minus 1.96 standard errors is about 9.4 points either side. That translation is ours and presumes the scale is anchored to a standard deviation of 15, which a reader should confirm for any product. The page on IQ standard deviation explains why that unit is conventional.
Stopping earlier is a choice with a price. Luijten and colleagues tested a rule based on standard error reduction, which stops when an extra item would improve the estimate by less than a set amount. In their data, responses from adolescents (mean age 13.7) to the Anxiety and Depressive Symptoms banks, the mean number of items fell from 9.98 and 8.13 under the default rules to 5.58 and 4.79 under their optimized rule. The authors report that for participants who reported no problems this meant fewer items but lower measurement accuracy and biased T-scores. That is a finding about patient-reported health measures, not intelligence tests, and it illustrates a general point: a shorter test through a stopping rule gives up some precision, and a vendor who says only "fewer questions" has not said how much.
5 How Many Items Does Adaptation Save? A Worked Example
When the bank holds an item near every level, an adaptive test can match a 40-item fixed form at the extremes of the scale with fewer than half the items, while near the middle the fixed form is as precise or better. The numbers below are our arithmetic from the standard 2PL information formula, not results from any study cited on this page, and they illustrate the mechanism rather than predict a real test. Assume every item has a discrimination of 1.5, an ability scale with a standard deviation of 1, and one scale unit equal to 15 IQ points. The fixed form has 40 items whose difficulties sit at the 40 midpoints of a normal distribution, so most items cluster near the average and a few sit at the extremes, which is how a form built for the general population looks. The adaptive test is idealized: it always finds an item exactly at the person's level.
Ability of the test taker
Fixed 40-item form: standard error (IQ points)
Adaptive, 18 on-target items
Adaptive, 40 on-target items
Average (theta 0, IQ 100)
0.25 (3.8)
0.31 (4.7)
0.21 (3.2)
One SD above (theta 1, IQ 115)
0.28 (4.2)
0.31 (4.7)
0.21 (3.2)
Two SD above (theta 2, IQ 130)
0.39 (5.9)
0.31 (4.7)
0.21 (3.2)
2.5 SD above (theta 2.5, IQ 137.5)
0.50 (7.5)
0.31 (4.7)
0.21 (3.2)
The table reads three ways. First, at the center of the scale the fixed 40-item form is more precise than 18 perfectly targeted items, 3.8 points against 4.7, because a form built for the general population puts most of its items there. Second, at IQ 130 the fixed form's 5.9 points is about a quarter larger than the adaptive test's 4.7, and at 137.5 its 7.5 points is about 60 percent larger, although the adaptive test used fewer than half as many items. Third, an adaptive test with the same 40 items would reach 3.2 points anywhere on the scale, which is the real prize: constant precision across the range, not merely brevity. The 18-item figure is what the arithmetic gives for a standard error of 0.32 (17.4 items, rounded up).
The same arithmetic shows why the bank matters. A standard error of 0.32 needs a total information of about 9.8, the reciprocal of 0.32 squared. Items on target supply 0.56 each, so 18 are enough. Items one standard deviation off target supply 0.34 each, and 30 are needed. Items two standard deviations off supply 0.10 each, and 97 are needed. A bank with no items near a person's level cannot adapt to that person, whatever the algorithm, and that is the working meaning of the phrase "a small bank cannot adapt": the freedom to choose is what produces the gain.
The idealization hides four costs. The opening items are chosen before the algorithm knows anything, so early information is lower. Maximum information selection favors the most discriminating items, which is the exposure problem taken up below, and exposure control deliberately gives some items that are not the single best choice. Shrinkage, noted in the ASVAB bulletin, pulls short-test estimates toward the mean. Treat the table as an upper bound on the benefit and a map of where it sits.
Published comparisons on real banks are more modest than the idealization. Choi and colleagues (2010) used post hoc simulations on the 28-item PROMIS depressive symptoms bank, comparing several static short-form strategies and the PROMIS CAT against the score from the whole bank. All the short forms and the CAT produced scores highly correlated with the full-bank score. The CAT outperformed each static short form in almost all criteria, but the short forms performed only marginally worse, and a two-stage branching format reduced the gap; the authors wrote that this semi-adaptive strategy came so close to CAT that it warranted further consideration. That bank measures a health symptom and is much smaller than the 602-word Toolbox bank, so the result cautions against assuming a large gain.
6 What Does Adaptive Testing Gain, and What Does It Lose?
Adaptive testing buys precision where a fixed form is weakest and, at equal precision, a shorter test, and it pays with a large calibrated bank, exposure risk, less freedom for the test taker and dependence on a statistical model. The Standards record the intended gains in the comment to Standard 4.3: common rationales are that score precision is increased, particularly for high-scoring and low-scoring examinees, or that comparable precision is achieved while testing time is reduced. The official ASVAB site states that the adaptive item selection process gives higher test-score precision and shorter test lengths than the paper-and-pencil ASVAB. The worked example explains the first claim: a fixed form is built for the middle, so its precision falls away at both ends.
The costs are documented just as clearly. The ITC guidelines list the potential advantages of adaptive designs as reduced testing time and cost, improved engagement, and increased precision and test security, and then the potential disadvantages: the inability of test takers to review previous answers, nonuniform item pool usage, and vulnerability of the adaptive algorithm to manipulation by test takers. The CAT-ASVAB bulletin gives an example of the third. Because a Bayesian estimate is drawn toward the mean of the prior, a low-ability applicant who answered only one or two items could obtain a score at or slightly below the mean, and the designers therefore built a penalty for incomplete tests.
There is also a cost in documentation. The Standards say, in the same comment, that adaptive tests are subject to the same requirements for documenting the validity of score interpretations for their intended use as other types of tests. The ITC guidelines add, in guideline 2.8, that the many design choices should be researched and documented, and the comment lists item bank size, IRT model, item selection method, exposure constraints, content constraints, scoring method and termination criteria.
Content coverage is a quieter cost. A fixed battery can guarantee that every person sees tasks from every ability it claims to measure, while selection on information alone can over-represent the most informative kind of item. For a broad battery spanning fluid reasoning, working memory and processing speed, adapting within one subtest is different from adapting across a profile, and the page on the CHC model lays out why those abilities are reported separately.
Adaptivity is a measurement engineering choice with a clear mechanism and clear costs, not a mark of quality. The page on reliability and validity separates two ideas that the label tends to blur: a precise estimate on a scale, and evidence that the scale means what the score report says it means.
7 What Is It Like to Take an Adaptive Test, and Does a Harder Test Mean a Higher Score?
An adaptive test is designed so that each person finds roughly half the items difficult, so the feeling of struggling says almost nothing about the score, and the share of correct answers says little either. In the NIH Toolbox design noted above, a strong and a weak candidate finish feeling about equally challenged, because the algorithm has moved each toward items at their own level. What counts is how hard the items answered correctly were, not how many were answered correctly, which is why the percentage correct on an adaptive test is a poor guide to the result.
Navigation differs across designs. In the ETS description of the GRE General Test, read on October 6, 2026, test takers can move forward and backward throughout an entire section, skip questions within it and go back to change answers, which fits a design that adapts between sections rather than within them. The ITC guidelines state the general point: multistage designs may allow test takers to review items within a stage, and the inability to review previous answers is a potential disadvantage of adaptive testing. The NIH Toolbox picture vocabulary test lets participants change their answer to the previous item only. Where a vendor says nothing about review, assume none is possible.
Timing interacts with adaptation too. The ETS page lists the Verbal Reasoning sections as 12 questions in 18 minutes and 15 questions in 23 minutes, for a whole test of about 1 hour 58 minutes since September 22, 2023. Time limits are a separate source of score variance from item difficulty, and the page on why IQ tests are timed explains what the clock measures.
A final point concerns repeat sittings. An adaptive test can draw different items on a retake, but whatever familiarity a person has gained with the item types and the format stays with them, and the page on how to prepare for an IQ test separates preparation that helps from familiarity that contaminates.
8 How Do Adaptive Tests Control Item Exposure, and How Large Must the Bank Be?
An adaptive test that always gave the single most informative item would give every person the same first items, so every practical system adds exposure control, and the control costs some precision. The CAT-ASVAB bulletin explains the problem: under a maximum-information rule, with every examinee assumed to start with equal ability, the first item would be the same for everyone and the second would be one of two choices, so the sequence is predictable and the initial items become overexposed. Georgiadou, Triantafillou and Economides (2007), reviewing exposure control strategies published from 1983 to 2005, add the mechanism by which exposure damages a test: examinees may become familiar with overexposed items and prepare for them, which lowers the items' actual difficulty, positively biases proficiency estimation and decreases the validity of the test. They also describe underexposed items that the algorithm rarely selects, which waste the cost of building a large bank.
The methods the two documents describe are a short list. The early CAT-ASVAB fix was the 5-4-3-2-1 procedure, in which the first item is chosen at random from the five most informative items, the second from the best four, and so on; the bulletin states that the net effect was still substantial use and overexposure of the most informative items. The randomesque strategy of Kingsbury and Zara (1989), as the review summarizes it, picks at random among a fixed number of the most informative items throughout the test. The Sympson-Hetter procedure of 1985, developed to meet the security requirements of the operational CAT-ASVAB, gives each item a parameter between zero and one, computed from repeated simulated administrations, that limits how often a selected item is actually used. The a-stratified approach of Chang and Ying (1999) starts from the argument that items with large discrimination values are chosen more often under maximum information selection, and groups items by discrimination so that all kinds are used.
Exposure control is why a real adaptive test does not always give the best item, and why the worked example is an upper bound. It is also why the bank must be large: the Standards say that in most cases large numbers of items are needed to construct a computerized adaptive test whose administered sets meet all the test specifications. The NIH Toolbox developers tried out 625 candidate words to end with a 602-word bank for a test of at most 25 items. A simple sum shows the limit: if a test administers 40 items from a bank of exactly 40, every person sees every item, and the test is a fixed form with extra software. The more the bank exceeds the test length, and the more evenly it spans the ability range, the more genuine choice the algorithm has.
Test security is the reason Mensa International gives for moving toward adaptive testing. Its FAQ page on getting your IQ tested, read on October 6, 2026, says paper tests may be coming to the end of their useful life, partly because many are old and partly because of the internet and the increasing likelihood that items are leaked online. It adds that since test takers on a computerized adaptive test receive different questions, there is no answer key that could be leaked as with fixed items, and that national groups will publish details. The page on the Mensa Norway test covers what is documented on the national side. The ITC guidelines name a limit to the argument: vulnerability of the adaptive algorithm to manipulation is listed among the disadvantages.
9 Which Real Tests Are Adaptive, and What Do Their Publishers Say?
Real adaptive instruments differ in what adapts, what stops the test and how much of the design is published, and the public documents of four institutional programs and four online tests or societies show the range. The table lists what each publisher states on the page named, as read on October 6, 2026 unless another date is given. Absences are reported as absences: a missing statement on a page means the page does not say, not that the design lacks the feature.
Instrument
What adapts
Selection and stopping, as published
Source and reading date
GRE General Test (ETS)
Section difficulty: the second Verbal and the second Quantitative section depend on the first
First section of average difficulty; second section depends on overall performance on the first; 12 then 15 questions per measure; skipping and review allowed within a section
ETS test structure page, October 6, 2026
CAT-ASVAB (U.S. Department of Defense)
Each item
Easier item after an error, harder after a correct answer; maximum information selection, 3PL model, Sympson-Hetter exposure control; fixed length (most subtests 15 scored questions)
Official ASVAB page, October 6, 2026; Technical Bulletin No. 1, March 2006
NIH Toolbox picture vocabulary and oral reading recognition
Each item
602-word bank, Rasch calibration; stops at a standard error below 0.3 or 25 items
Gershon and colleagues, 2014
PROMIS CATs (HealthMeasures)
Each item
Items chosen from a bank by previous answers; 4 to 12 items; default stop at a standard error of 0.32 in data analyzed by Luijten and colleagues
HealthMeasures page, October 6, 2026; Luijten and colleagues, 2025
CAT-II (Cognitive Metrics)
Each item
Page says it uses IRT to maximize information from each response; 30 to 60 items; ends when accuracy reaches at least 0.925; about 15 to 30 minutes; 10 dollars for the full score profile; bank size and exposure control not stated on the page
cognitivemetrics.com/test/CAT, October 6, 2026
BRGHT
Difficulty level
40 questions in 30 minutes; 20 difficulty levels; one level up per correct answer and one down per error; score is points equal to level; Rasch and 2PL models from over 500,000 test takers; page does not address item exposure or reliability
brght.org/method, October 6, 2026
JCTI (cogn-iq.org)
Each item
Adaptive form since 2025 with a mean of 29.4 items and a range of 20 to 42, per the manual as recorded in ACIS's own review; the manual site could not be opened by our fetch tool on October 6, 2026
ACIS review of Cognitive Metrics, read September 20, 2026
Mensa International
Announced
Says it is moving to computerized adaptive testing because of leaked items; no selection rule, stopping rule or bank size given on the page
Mensa FAQ page, October 6, 2026
Three observations follow. First, the programs that document the most are the ones whose users must defend scores: the ASVAB bulletin and the Toolbox paper publish models, stopping rules and exposure methods, the Toolbox paper states its bank size, and the PROMIS item range is on the HealthMeasures page while its stopping defaults appear in research papers. Second, the online products vary. CAT-II states a selection principle, an item range and a stopping threshold, but the page gives no bank size or exposure control, and its "accuracy" figure is not defined there, so we do not translate it into a reliability. BRGHT's method page is explicit that its rule is a one-step staircase rather than model-based selection, and it does not discuss exposure or reliability. The pages on the BRGHT adaptive, self-normed design and on Cognitive Metrics work through those products.
Third, the label hides a spectrum. A section-level design such as the GRE keeps review within each section. A staircase gives rough targeting without weighing items by discrimination. A model-based item-level design such as CAT-ASVAB extracts the most information per item and needs the most infrastructure. The pages on how the GMAT adapts to a score and the SAT comparison treat admission exams on their own terms.
Individually administered clinical batteries, such as those described on the WAIS-5 subtests and Woodcock-Johnson V pages, are given by a trained examiner, and their publishers describe their administration rules in their manuals. A reader should not assume that "adaptive" in the item response sense applies to them without checking each manual. Research platforms raise a different question, which tests suit an unsupervised sample, and the page on measuring IQ in online studies covers it.
10 Does an Adaptive Label Say Anything About Norms, Accuracy or Validity?
It does not: adaptive describes how items are delivered and scored, while the meaning of an IQ score depends on a reference group and on evidence that the scale measures what it claims. An adaptive algorithm can estimate an ability on its own scale very precisely and still leave the key questions open. Precise relative to what? The scale of an IRT model is arbitrary until someone fixes it, and an IQ-style score requires a norm sample, which is the subject of the page on how IQ scores are normed. BRGHT's method page, for example, says the raw score is compared to the entire population of its test takers to calculate an IQ with a median of 100 and a standard deviation of 15. That is a statement about the reference group, not about adaptivity, and a reader comparing it with a test normed on a representative sample is comparing two different things under the same label.
Precision and validity are also different properties. The standard error that stops a test, such as 0.3 for the NIH Toolbox, describes how stable the estimate is on that test's own scale; it says nothing about whether the items measure vocabulary, reasoning or familiarity with test formats, and the page on what IQ scores mean separates the score from the construct. The Standards say in the comment to Standard 4.3 that adaptive tests carry the same validity documentation requirements as other tests, and the ITC guidelines add in guideline 1.41 that, when adaptive technology-based assessments are used, reported information should include what the test measures for specific individuals or groups of test takers.
Breadth is a third issue. An adaptive test over one bank measures one thing: the Toolbox language measures are two separate tests that give two scores. A Full Scale IQ and a profile of indices require tasks from several abilities, and an adaptive test can make one scale efficient without turning one ability into six.
For a reader weighing a free adaptive test, the practical sequence is to separate the questions: what rule adapts, who the reference group is, what error is reported, and what evidence supports the intended use. The pages on whether online IQ tests are accurate and on how to tell if a score is real work through the last three for online products in general.
11 How Can You Check Whether a Test Called Adaptive Really Adapts?
Five questions, each answerable from a method page or manual, separate a model-based adaptive test from one that borrows the word, and a test that cannot answer them has told you what it documents. The questions come from the documentation the Standards and the ITC guidelines ask of adaptive tests. Standard 4.3 says developers should document the rationale and supporting evidence for the administration, scoring and reporting rules used in tests that use computer algorithms to select items, and that this documentation should include the procedures used in selecting items or sets of items, in determining the starting point and termination conditions, in scoring the test, and in controlling item exposure. The ITC guidelines list the same components: a large and fully calibrated bank, a starting point, a selection algorithm, a scoring algorithm and termination criteria.
Question
What a complete answer looks like
Example from this page
1. What picks the next item?
A named rule that uses your answers: maximum information from a stated model, a section-level branch, or a fixed step rule
CAT-ASVAB uses maximum information; GRE branches by section; BRGHT steps one level
2. What ends the test?
A fixed count, a standard error target, a time limit, or a stated combination with minimum and maximum lengths
CAT-ASVAB fixed length; NIH Toolbox standard error below 0.3 or 25 items
3. What is behind the questions?
A bank much larger than the test, a stated model, a stated calibration sample and a stated exposure method
NIH Toolbox 602 words from 625 tried out; CAT-ASVAB Sympson-Hetter
4. How is the score estimated, and with what error?
A named estimator and a standard error or interval that can differ by person
Bayesian mode in CAT-ASVAB; standard error of theta in PROMIS
5. Compared with whom, and for what use?
A described reference group and evidence for the use the score is put to
The Standards require validity documentation for adaptive tests as for others
One further check needs no manual. Note whether the questions get noticeably easier after you miss one and harder after you answer one correctly, which a model-based item-level test will show and a section-level test will not show until the next section begins. The check proves nothing about validity, but it tells you whether the word "adaptive" describes the experience.
12 Where ACIS Sits, and What a Buyer Should Do
ACIS is not an adaptive test: for the form you buy, every person gets the same fixed set of subtests, and the cost of that choice is time, which a buyer should weigh against what the score will be used for. The facts below were read on the ACIS home page on October 6, 2026, and prices can change. The three forms are one time purchases with no subscription: Quick at 15 dollars with 6 subtests in 3 domains and about 45 minutes, Optimized at 30 dollars with 13 subtests in 5 domains and about 110 minutes, and Full Scale at 50 dollars with all 20 subtests in six domains and about 175 minutes. Breaks are allowed, there is a free trial with no card, a 5 day quality guarantee and 30 days to complete a form. The report gives a Full Scale IQ and six primary indices on the standard scale with a mean of 100 and a standard deviation of 15, scaled subtest scores from 1 to 19, percentiles and a 95 percent confidence interval, with adult norms covering ages 16 to 90. The technical manual documents the instrument, and this page quotes no figure from it.
The limits are plain. ACIS is online and unsupervised, it is not a clinical or diagnostic instrument, it is not for hiring, school accommodations or admission to high IQ societies, and it is available in English only. A fixed form asks everyone the same tasks, so it cannot save time by targeting items, and a 175 minute battery takes far longer than a Toolbox vocabulary test that typically takes five minutes or less. In exchange, it measures six domains in a single battery, gives every buyer of a form the same set of subtests, and shows an interval on the score report. An adaptive design can offer shorter administration and, in the idealized arithmetic above, tighter precision at the extremes than a fixed form built for the general population, and a buyer who needs only one narrow ability quickly may be better served by a documented adaptive instrument. This page makes no claim about the precision of ACIS beyond the interval its report shows.
What a buyer should do is apply the five questions to every product, including this one. The Standards for Educational and Psychological Testing (AERA, APA and NCME, 2014), which the American Psychological Association co-publishes with the other two bodies, give the reading rule for a score: it comes with error, and the conditional standard error of measurement is the error at a given score level, so a relatively large error means relatively low precision at that level. For an adaptive test that error varies by person, so ask whether the reported error is specific to you. An adaptive format does not change the limits on use: decisions about diagnosis, accommodations or employment need an appropriate instrument, administered by a qualified professional, with evidence for that purpose. The pages on types of IQ tests and the IQ test for adults set out the options, and the page on how long an IQ test takes lists times by instrument.
Every product fact and figure on this page comes from the documents below, read or verified on October 6, 2026, and figures we derived ourselves are labeled in the text as our arithmetic.
American Educational Research Association, American Psychological Association, and National Council on Measurement in Education (2014). Standards for Educational and Psychological Testing. Washington, DC: AERA. Open access edition: https://www.testingstandards.net/open-access-files.html
Gershon, R. C., Cook, K. F., Mungas, D., Manly, J. J., Slotkin, J., Beaumont, J. L., and Weintraub, S. (2014). Language measures of the NIH Toolbox Cognition Battery. Journal of the International Neuropsychological Society, 20(6), 642 to 651. https://doi.org/10.1017/S1355617714000411
Luijten, M. A. J., Schalet, B. D., Roorda, L. D., Haverman, L., and Terwee, C. B. (2025). Reducing patient burden of PROMs in healthcare through advanced computerized adaptive testing stopping rules. Quality of Life Research, 34(11), 3205 to 3214. https://doi.org/10.1007/s11136-025-04079-7
Choi, S. W., Reise, S. P., Pilkonis, P. A., Hays, R. D., and Cella, D. (2010). Efficiency of static and computer adaptive short forms compared to full-length measures of depressive symptoms. Quality of Life Research, 19(1), 125 to 136. https://doi.org/10.1007/s11136-009-9560-5
Georgiadou, E., Triantafillou, E., and Economides, A. A. (2007). A review of item exposure control strategies for computerized adaptive testing developed from 1983 to 2005. Journal of Technology, Learning, and Assessment, 5(8). https://files.eric.ed.gov/fulltext/EJ838610.pdf
Lord, F. M. (1980). Applications of Item Response Theory to Practical Testing Problems. Erlbaum (reissued 2012 by Routledge). https://doi.org/10.4324/9780203056615
Wainer, H., Dorans, N. J., Flaugher, R., Green, B. F., and Mislevy, R. J. (2000). Computerized Adaptive Testing: A Primer (2nd ed.). Mahwah, NJ: Erlbaum. https://doi.org/10.4324/9781410605931
An adaptive IQ test chooses each question from your earlier answers, giving harder items after successes and easier ones after errors, using a model of item difficulty. The aim is a precise estimate from fewer items. The label describes how items are delivered, so a score still needs a reference group.
How does a computerized adaptive test work?
It keeps a running ability estimate, picks the unused item expected to be most informative at that estimate, updates the estimate after each response, and stops when a rule is met. The score comes from the pattern of responses and the calibrated item parameters, not from a raw count of correct answers.
Is an adaptive IQ test more accurate than a regular one?
Not automatically. Adaptivity can raise precision at very high and very low ability and can shorten a test, but accuracy also depends on the quality of the item bank, the norms and the stopping rule. Where researchers compared real banks, well chosen short fixed forms came close to the adaptive result.
How many questions does an adaptive IQ test have?
There is no standard number. It depends on the stopping rule and the bank: a fixed-length rule gives everyone the same count, while a precision rule gives each person as many items as needed, usually with a minimum and a maximum. Always check whether a published count is a cap, an average or a fixed length.
Why does an adaptive test feel harder as it goes?
Correct answers lead to harder items. The algorithm steers each person toward items they have about an even chance of answering correctly, so almost everyone feels challenged. That feeling reflects the selection rule rather than how well you are doing, and the share of correct answers is a poor guide to your score.
Can you go back to a previous question in an adaptive test?
Usually not in item-level designs, because later items depend on earlier answers, and the International Test Commission lists inability to review as a disadvantage. Section-level designs differ: the GRE lets you move back and forth within a section. Check the instructions of the specific test before you begin.
What does item response theory have to do with adaptive testing?
Item response theory supplies the model that places items and people on one scale, which makes scores comparable when people saw different items. It also gives each item an information value at each ability level, which the algorithm uses to choose the next question. Without calibrated item parameters, a test cannot select or score adaptively.
Is the GRE an adaptive test?
Yes, at the section level. ETS says the first Verbal and Quantitative section is of average difficulty and the difficulty of the second depends on performance on the first. Within a section the test does not adapt question by question, and test takers can skip questions and return to them.
Is the ASVAB adaptive?
The computerized CAT-ASVAB is. The official site says an incorrect answer leads to an easier item and a correct answer to a harder one, drawn from a pool ranging from very easy to very hard. Its technical bulletin describes maximum information selection, fixed test lengths and probabilistic exposure control.
What is a stopping rule in adaptive testing?
A stopping rule is the condition that ends the test. It can be a fixed number of items, a target standard error such as 0.3, a time limit, or a combination with minimum and maximum lengths. The choice decides whether people receive equal length or equal precision, and the two cannot both be guaranteed.
What is item exposure control?
Exposure control limits how often any one item is given, because always choosing the most informative item would overuse a few. Methods include randomizing among the best items and probabilistic limits such as Sympson-Hetter. Overexposed items can become easier through familiarity, which biases scores upward and reduces validity.
Do the NIH Toolbox and PROMIS use adaptive testing?
Yes. The NIH Toolbox picture vocabulary and oral reading recognition tests are computer adaptive and scored with item response theory, and PROMIS offers computer adaptive tests that administer between 4 and 12 items. Both belong to health and neurological measurement programs rather than consumer IQ testing.
Does Mensa use adaptive testing?
Mensa International says it is moving to computerized adaptive testing because paper tests are aging and items can be leaked online, and that national groups will publish details. The page we read on October 6, 2026 gave no selection rule, stopping rule or bank size, so check your national group.
How large does an adaptive test item bank need to be?
No single number applies. The Standards say large numbers of items are generally needed, and the International Test Commission lists bank size among the design choices to document. The bank needs items near every ability level, so a bank only slightly larger than the test cannot adapt in any meaningful way.
Does a harder adaptive test mean a higher IQ?
No. Difficulty feels similar to everyone because items are targeted near each person's level. What matters is which items were answered correctly and their calibrated difficulty. Judge a test by its published scoring rule and its norms, not by how hard it felt while you were taking it.
Should I trust a free adaptive IQ test?
Trust it as far as it documents its design and its norms. Look for a named selection rule, a stopping rule, the bank calibration, a reported standard error and a description of the reference group. Adaptive delivery alone does not make a score meaningful or comparable with scores from other tests.
How can I tell whether a test that calls itself adaptive really adapts?
Look for a selection rule that depends on your answers, a stopping rule, a bank larger than the test, an estimation method and a reported error. If two people can see identical questions in identical order regardless of their answers, the test is fixed, whatever its marketing calls it.
Is ACIS an adaptive test?
No. ACIS gives every person the same fixed set of subtests for the form they buy, and its report gives a Full Scale IQ, six indices, percentiles and a 95 percent confidence interval. The technical manual documents the instrument, and the ACIS pages say plainly that it does not adapt.
Do adaptive tests report a margin of error?
They can, and good ones do, because a standard error comes with every ability estimate. It may differ by person, since precision depends on how well the bank covers your level, so ask for a person-specific or conditional error rather than a single average figure for the whole test.
Can an adaptive IQ test be used for hiring, school placement or diagnosis?
An adaptive format does not change use limits. Any use needs evidence for that specific purpose, and decisions about diagnosis or accommodations need a qualified professional and an appropriate instrument. An unsupervised online test, adaptive or fixed, is not a substitute for that process.
Is a shorter adaptive test better than a longer fixed battery?
It depends on the question. A short adaptive test can estimate one ability efficiently, while a longer fixed battery measures several abilities and reports a profile. Choose by what the score will be used for and what the publisher documents, not by the number of minutes.
Take the assessment
You get a profile, not a number
ACIS measures six CHC domains across 20 subtests and reports each one with its own normed score and confidence interval, so you can see where you are strong and where you are not.