Test Length

The quick IQ test and what it costs

A short intelligence test is not automatically a bad one. Some of the best documented instruments in the field run under half an hour. What separates them from a five minute web quiz is whether the test was built short and normed short, or whether it is a long battery with most of its items taken out. This page works the arithmetic and names which is which.

Two hands in a blue sleeve holding a black twin bell alarm clock whose face reads 2026, with the hands near twelve.
Minutes and measurement error are the same conversation: the Spearman-Brown formula, published in 1910, predicts exactly what a shorter test costs.

0 Quick Answer

A quick IQ test can produce a usable estimate, but only of a narrow slice of ability, and only if the short form was standardized as a short form. The cost of shortening is not a matter of opinion. It is set by a formula published independently by Charles Spearman and William Brown in volume three of the British Journal of Psychology in 1910, and it converts directly into the width of the band around your score. Cut a test's length and reliability falls. Reliability falling widens the standard error of measurement. A wider standard error means the same reported number covers more ground.

Real short instruments exist and deserve to be taken seriously. The two subtest form of the Wechsler Abbreviated Scale of Intelligence runs about 15 minutes. The Kaufman Brief Intelligence Test, Second Edition runs about 20. The National Adult Reading Test is 50 words read aloud. All three were built as short instruments, standardized as short instruments, and published with their limits attached. That is a different category from a free page offering a five minute IQ test with no reference sample behind it.

The duration question in general is covered on the page on how long an IQ test takes, per item timing on the page on time limits, and the free tier on the free test page. This page owns the narrower question underneath all three: what precision costs, and what the shortest defensible option looks like once you have paid it.

2.40 to 5.22

Range of published ACIS index standard errors of measurement, from verbal comprehension to processing speed.

.87 and .92

Correlations of the WASI two and four subtest Full Scale IQs with the WAIS-III Full Scale IQ, reported by Axelrod in Assessment in 2002.

7.68 points

Standard error of estimate for a Full Scale IQ predicted from the 50 word National Adult Reading Test, Crawford and colleagues, 1989.

1 The Arithmetic That Sets the Error Bar

There is exactly one equation between a test's reliability and the width of the band around your score, and it has no free parameters to argue about. The standard error of measurement equals the standard deviation of the score scale multiplied by the square root of one minus the reliability. IQ scores use a standard deviation of 15, so the whole calculation is 15 times the square root of one minus r. Run it backward and reliability equals one minus the square of the standard error divided by 15.

That inversion is worth doing on a real instrument, because it shows the relationship is arithmetic rather than assertion. The table below takes the nine ACIS composite standard errors published in the technical manual, recovers the reliability each one implies, and then converts it into the width of a 95 percent interval at 1.96 standard errors on each side. Compare the recovered column against the published omega column: they agree to the third decimal in every row, which is what it looks like when a publisher's numbers are internally consistent.

ScoreIndicatorsPublished omegaPublished SEMReliability recovered from the SEMWidth of the 95 percent interval
GAI15.98851.61.98856.3 points
CFI15.98371.92.98367.5 points
SAI9.98231.99.98247.8 points
VCI5.97452.40.97449.4 points
FRI5.97272.48.97279.7 points
VSI3.94883.39.948913.3 points
QRI2.92834.02.928215.8 points
WMI3.92474.12.924616.2 points
PSI2.87905.22.878920.5 points

Read the indicators column against the interval column and the pattern is unmistakable. Fifteen indicators buy a band about six points wide. Five indicators buy about nine. Two buy sixteen to twenty. Nothing else changed between those rows: same scale, same reference frame, same test taker. The only difference is how much measurement was performed before a number was printed.

The metric itself, and why 15 is the standard deviation everyone uses, is explained on the page on the 15 point standard deviation. What the resulting number means in ordinary classification language is on the IQ score scale.

2 What the 1910 Formula Predicts When You Drop Indicators

The Spearman-Brown prediction formula turns the previous table into a forecast, and applying it to a real index shows what each subtest you remove actually costs. Spearman and Brown published it independently in the same 1910 volume of the British Journal of Psychology. Spearman's paper was Correlation calculated with faulty data, pages 271 to 295. Brown's was Some experimental results in the correlation of mental abilities, pages 296 to 322. The formula says that if you multiply a test's length by a factor k, the new reliability equals k times r divided by one plus k minus one times r.

Its inverse is the useful direction here. Given a composite reliability built on five indicators, you can recover the reliability a single indicator would carry, then project back up to any other number of indicators. Take the ACIS verbal comprehension index, published at an omega of .9745 across five subtests. The formula puts the implied single indicator reliability at about .884. From there the projection runs as follows.

Indicators in the indexProjected reliabilityProjected SEMProjected 95 percent widthWhat it means at 132
5, the published Full Scale form.97452.409.4 points127 to 137
4.9682.6710.5 points127 to 137
3.9583.0712.0 points126 to 138
2.9393.7214.6 points125 to 139
1.8845.1020.0 points122 to 142
What these figures are and are notOnly the first row is a published ACIS figure. Rows two through five are projections produced by applying the 1910 formula to that published value, and they carry the formula's assumption that every subtest you remove was as good as the ones you kept. Real removals are rarely that tidy. Treat the projected rows as a best case for what shortening costs, not as a measured result, and read the published values themselves in the manual.

Two things fall out of that table. The first is that the cost of the fifth subtest is small: going from five indicators to four widens the band by about one point. The second is that the cost accelerates. Going from two indicators to one widens it by five and a half. Shortening is cheap at the long end and expensive at the short end, which is why abbreviated forms of good instruments are defensible for screening and why five item quizzes are not defensible for anything. The general derivation, applied to whole battery durations rather than to index indicators, is worked through on the duration page.

One caution about the formula's reputation. It predicts what happens when you add or remove parallel items of equal quality. It says nothing about what happens when you remove a whole domain, because that is not a change in length at all. It is a change in what the test measures, which is the subject of the section on built short against cut short below.

3 What a Confidence Interval Means When You Are Near a Line

Almost nobody wants a score for its own sake. They want to know which side of a line they are on, and that is the question a wide interval destroys. Suppose two people are told 132. One took a form with a standard error of 2.40 points. The other took a quiz with a reliability of .70, which puts its standard error at 8.2 points. Both are looking at the same three digits on the same scale, and they are in completely different epistemic positions.

Put a threshold at 130 and work the probabilities. Treating the reported score as the best available estimate of the true score, the chance that the true score exceeds 130 is the area of a normal curve above the threshold, with the standard error as its spread. At a standard error of 2.40, the threshold sits 0.83 standard errors below the reported score, giving roughly an 80 percent chance of clearing it. At 3.72, the projected two indicator figure from the previous section, the gap is 0.54 standard errors and the chance drops to about 70 percent. At 8.2 it is 0.24 standard errors and about 60 percent, which is barely better than a coin weighted slightly in your favor.

About 80 percent

Chance a reported 132 sits above 130 at a standard error of 2.40 points.

About 70 percent

The same reported 132 at a standard error of 3.72 points.

About 60 percent

The same reported 132 at a standard error of 8.2 points, typical of a short homogeneous quiz.

The correction that makes the short test look worse stillThe arithmetic above treats your obtained score as your best estimate of your true score, which is the convention most reports use. Classical test theory says something less flattering: the best estimate of a true score regresses toward the mean in proportion to the reliability. At a reliability of .9745, a reported 132 gives an estimated true score of about 131. At .70 it gives about 122. The low reliability case does not merely widen the band around 132, it moves the center of the band down, and almost no consumer test says so.

This is why a responsible report prints an interval rather than a point, and why comparing two scores from two different tests without their intervals is not a comparison at all. Converting a figure into a rank is handled on the percentile calculator, and what the classification bands mean once the interval is known is on the page on what counts as a good score.

There is a practical corollary for anyone testing near a threshold for a specific reason. If the number has to clear a line, the only sensible response to a wide interval is more measurement, not a second short test. Which instruments actually sit at each end of that trade is the subject of the next four sections.

4 Built Short Against Cut Short

The single most useful distinction in this whole subject is whether a short test was designed and standardized as a short test, or assembled by deleting items from a long one. The first is a legitimate instrument with its own norms and its own error terms. The second inherits a reputation it has not earned, because a parent form's validity evidence does not transfer to a subset of its items without new evidence.

That argument has a canonical citation. Gregory T. Smith, Denis M. McCarthy, and Kristen G. Anderson published On the sins of short-form development in Psychological Assessment, volume 12, issue 1, pages 102 to 111, in 2000. They argued that the empirical short form literature has been characterized by overly optimistic views of the transfer of validity from the parent form to the short form, and by weak application of psychometric principles in validating short forms. They then set out two general and nine specific methodological failures, and argued the validity standards for short forms should be high rather than relaxed.

The reason the transfer fails is mechanical. Deleting alternate items from a subtest increases the difficulty gap between adjacent items, which changes the test taking experience and makes it harder to locate where a person's performance starts to break down. It changes the order in which tasks are encountered, which affects warmup and fatigue. And it changes the content sampled, so a composite that was built to average across several kinds of task ends up averaging across fewer.

  • Built short. New items written and calibrated for a short instrument, a dedicated standardization sample, published reliabilities computed on the short form itself, and error terms reported for the scores the short form actually produces.
  • Cut short. A subset of a longer test's items, scored against the long test's tables, with validity claimed by inheritance rather than demonstrated.
  • Neither. Items written from scratch with no reference sample at all, scored by an unpublished rule. This is where most free five minute pages sit, and it is not a short form of anything.

Test publishers know this, which is why the good short instruments carry their own manuals rather than an appendix in somebody else's. The broader map of which instruments belong to which family is on the page on types of IQ test, and the consumer version of the same distinction is on the free against validated comparison.

5 The Best Documented Short Form: the WASI Family

If you want to see what a properly built short intelligence test looks like, the Wechsler Abbreviated Scale of Intelligence is the reference case, and its published figures are unusually specific. The original WASI was published in 1999 and normed on a nationally representative sample of 2,245 individuals. Its two subtest Full Scale IQ carried reliability coefficients of .93 for children and .96 for adults, and its four subtest form carried .96 for children and .98 for adults, figures reported by Eric Pierson, Lydia Kilmer, Barbara Rothlisberg, and David McIntosh in the Journal of Psychoeducational Assessment, volume 30, issue 1, pages 10 to 24, in 2012.

The validity evidence comes from outside the publisher. Bradley Axelrod compared short form estimates against full administrations in a mixed clinical sample of 72 participants and published the results in Assessment, volume 9, issue 1, pages 17 to 23, in 2002. The four subtest WASI Full Scale IQ correlated .92 with the WAIS-III Full Scale IQ, and the two subtest version correlated .87. Those are the numbers a well built short form earns against its parent battery.

The second edition arrived in 2011 with a norming sample of 2,300 people aged 6 to 90, split as 1,100 children and 1,200 adults. Its manual reports internal consistency from .87 to .92 for the four subtests, .94 to .95 for the index scores, and .96 to .97 for the Full Scale IQ. Pearson lists the WASI-II at 30 minutes for the four subtest form, 15 minutes for the two subtest form, a qualification of Level C, and a kit price of 530.40 dollars.

FormSubtestsTimePublished reliabilityCorrelation with the full battery
WASI FSIQ-4, 1999Vocabulary, Similarities, Block Design, Matrix ReasoningAbout 30 minutes.96 children, .98 adults.92 with WAIS-III FSIQ, Axelrod 2002
WASI FSIQ-2, 1999Vocabulary, Matrix ReasoningAbout 15 minutes.93 children, .96 adults.87 with WAIS-III FSIQ, Axelrod 2002
WASI-II, 2011Same four subtests, new and extended items15 or 30 minutes.96 to .97 for the FSIQLinked to WAIS-IV in samples of 182 and to WISC-IV in samples of 201

Notice what a 15 minute test buys when it is done properly. A correlation of .87 with a full battery means the short form and the long form share about three quarters of their variance, and leave a quarter unaccounted for. That is a strong result and it is still a quarter. Pearson's own description of the WASI-II is honest about the intent: the product page frames it as a way to "Screen to determine if in-depth intellectual assessment is needed."

6 What Even the WASI-II Cannot Do

A well built short form measures general ability well and almost nothing else, and the evidence for that comes from reanalyzing the publisher's own correlation matrices. Brittany McGeehan, Nadine Ndip, and Ryan J. McGill published an exploratory bifactor and Schmid-Leiman analysis of the WASI-II in Archives of Assessment Psychology, volume 7, number 1, pages 7 to 27, in 2017. Their point of departure was that the WASI-II manual implies a hierarchical structure but never tested one.

The results are stark. In the adult sample of 1,200, the general factor accounted for 61.2 percent of the total variance and 75.7 percent of the common variance. Omega hierarchical for general intelligence came out at .84, comfortably high enough to interpret. Omega for the two group factors, Verbal and Perceptual, came out at .16 each. The authors concluded that those two specific factors "likely possess too little true score variance for confidant clinical interpretation," and recommended that users focus most, if not all, of their interpretive weight on the Full Scale composite.

Their sharper sentence is worth quoting because it applies directly to any short test that prints multiple index scores: "Given the lack of target construct specificity in the VC and PR index scores, it is our position that their provision on the WASI-II is an invitation for misuse."

The general rule this impliesTwo indicators can estimate a general factor. Two indicators cannot estimate a specific factor that is distinct from the general factor, because after g has taken its share there is not enough reliable variance left. This is why a short test can tell you roughly where you sit overall and cannot tell you that your verbal ability outruns your spatial ability. A profile needs indicators per domain, not indicators in total.

The same logic explains why the ACIS index reliabilities in the first table fall as indicators fall, and why the general ability composite built from 15 subtests reaches an omega of .9885 while a two subtest index reaches .8790. What a general factor is, and why it dominates every cognitive battery ever built, is explained on the g factor page. What the separate domains add on top of it is on the cognitive domains page.

7 Three More Short Instruments Worth Knowing

The short test market did not begin with the WASI, and three older instruments show three different strategies for getting an estimate fast. Each is legitimate within its own claims, and each has a limit that follows directly from its design.

InstrumentPublishedWhat it consists ofTimeDocumented relationship to a full battery
Ammons Quick TestAmmons and Ammons, Psychological Reports, 11, 111 to 162, 1962A passive response picture vocabulary task, the examinee points at one of several drawingsMinutesValidity coefficients reported across a very wide band in the follow up literature, which is itself the finding
National Adult Reading TestNelson, 1982; Nelson and Willison, 1991Fifty short words of irregular pronunciation, read aloud and scored for errorsTwo to three minutesPredicted 66 percent of WAIS Full Scale IQ variance in a cross validation of 151, Crawford and colleagues, 1989
KBIT-2Kaufman and Kaufman, 2004Verbal Knowledge with 60 items, Riddles with 48, and Matrices with 46About 20 minutesStandardized on 2,120 individuals, positioned as a screener for ages 4 to 90

The NART is the most instructive of the three because its arithmetic is fully public. John Crawford, David Parker, Lorraine Stewart, Jean Besson, and Gerard De Lacey reported in the British Journal of Clinical Psychology, volume 28, pages 267 to 273, in 1989, that Nelson's regression equations predicted 66, 72, and 33 percent of the variance in WAIS Full Scale, Verbal, and Performance IQ in a cross validation sample of 151 adults aged 16 to 88. Combining the standardization and cross validation samples to 271 people, their new equation for Full Scale IQ carried a standard error of estimate of 7.68 points.

Sit with that last figure. A three minute reading task predicts Full Scale IQ with a standard error of 7.68 points, which puts a 95 percent interval about 30 points wide around the estimate. That is a genuinely useful clinical tool, because a neuropsychologist comparing a current score against a premorbid estimate can work with a 30 point band. It is close to useless as a statement about where an individual sits on the scale, which is a different job the instrument never claimed.

The KBIT-2 shows the third strategy: keep the task types but keep only one indicator each. Jason Skues and Everarda Cunningham described it in the Australian Journal of Educational and Developmental Psychology, volume 13, pages 16 to 27, in 2013, as a brief individually administered test taking 15 to 30 minutes that can be administered by non psychologists. Pearson lists it at about 20 minutes, ages 4 to 90, qualification Level B, and describes its purpose as identifying people who "require a more comprehensive evaluation."

8 What the Free Five Minute Pages Actually Are

The free pages that rank for five minute IQ test are not short forms of anything, and treating them as a cheap version of the instruments above is the mistake this page exists to prevent. They are their own category, and the honest description of that category is short: an item set with a scoring rule and no documented reference sample.

Three examples read in September 2026 give the shape of it. One site offers a 12 question quick quiz of about five minutes and a 25 question full quiz of about ten, both free, and states in its own disclaimer that its results "are not equivalent to professionally administered cognitive assessments." Another runs 30 questions in 20 minutes and issues an instant certificate, describing no standardization sample anywhere this article could find. A third publishes a 20 item paper folding task under Alfred Binet's name and states outright that free tests of that kind do not provide professional assessments of any kind.

Two of those three are more honest about themselves than a great deal of paid competition. The problem is not deception in every case. The problem is that a raw score becomes an IQ only by comparison against a documented group of known composition and known age structure, and none of these pages has one. Without it, the number on the screen is a rank within an unspecified population, which is not a rank at all. How that conversion is built is set out on the page on norming.

What a five minute test is legitimately good forCuriosity, practice at unfamiliar item formats, and the pleasure of solving puzzles. None of that requires an apology. What it cannot do is support a decision, settle an argument, or justify a claim to anybody else, and a certificate at the end changes none of that. If a five minute result is reported to a decimal place, the precision is decorative. The wider comparison is on the page on online test accuracy and the page on what paying actually buys.

There is one more asymmetry worth naming. A short test that underestimates you produces annoyance and a retake. A short test that overestimates you produces a belief, and beliefs are stickier than annoyance. The flattering error is the one that survives, which is why the interval matters most in exactly the direction people are least inclined to check.

9 What Five Minutes Can and Cannot Measure

Five minutes is not uniformly useless. It is useless for some things and nearly adequate for one, and the difference is whether the ability in question is measured by accuracy or by rate. Power tasks, where items get harder and the question is how far you get, need many items across a difficulty range to place a person. Speeded tasks, where items are easy and the question is how many you finish, accumulate information fast because every second produces data.

That is why processing speed subtests are short by design in every major battery and why they are not short by compromise. ACIS reports processing speed from two subtests, Symbol Search and Coding, and the index carries a published standard error of 5.22 points against 2.40 for verbal comprehension. It is the least precise index on the scale and the least g loaded at .648, and it is still worth reporting because it carries information the other five do not. What that domain actually captures is on the processing speed page.

Reasoning is the opposite case. A fluid reasoning item can take a minute to solve on its own, so a five minute window holds perhaps five items, and five items cannot span a difficulty range wide enough to place anyone. A vocabulary estimate degrades more gracefully, because vocabulary items are quick and highly correlated with each other, which is exactly why so many brief instruments lean on receptive vocabulary. The Ammons Quick Test, the NART, and the KBIT-2 Verbal Knowledge subtest all take that route.

.648

Published g loading of the ACIS processing speed index, the lowest of the six primary indices.

.864 and .922

Published g loadings of verbal comprehension and fluid reasoning, the two domains a short test compromises most.

.475

Published correlation between the working memory and processing speed indices, the weakest pair on the scale.

That last figure is the argument against reading any one short measure as a stand in for the whole. Working memory and processing speed correlate at .475 on the ACIS scale, so a person can be strong on one and ordinary on the other without anything being wrong. A five minute test samples one of them, at best, and reports it as though it were the answer. The full set of index relationships is in the technical manual, and the framework organizing them is on the CHC model page.

10 The Shortest ACIS Form and Exactly How It Is Framed

ACIS sells a short form, this page is published by ACIS, and the honest way to handle that is to state what the short form gives up in the same paragraph that names its price. The Quick form costs 15 dollars, runs six subtests in about 45 minutes, and covers three of the six domains: verbal comprehension, fluid reasoning, and working memory. The six subtests are Similarities and Vocabulary for verbal comprehension, Matrix Reasoning and Figure Weights for fluid reasoning, and Digit Span and Alphanumeric Sequencing for working memory.

The ACIS technical manual describes it in its own words as a focused estimate of selected verbal, fluid, and working memory ability, states that it "can support a useful screening level profile," and states that it is not the same as a 20 subtest Full Scale administration. It also records that the three forms "are not merely shorter or longer versions of the same score" and that a report must name which form was completed, because the form determines which scores can be produced at all.

Apply the Spearman-Brown projection from earlier to see what that means numerically. Verbal comprehension is published at five indicators. Quick supplies two. Fluid reasoning is published at five. Quick supplies two. Working memory is published at three. Quick supplies two.

IndexFull Scale indicatorsPublished SEMQuick indicatorsProjected SEM at two indicatorsProjected 95 percent width
VCI52.402About 3.72About 14.6 points
FRI52.482About 3.84About 15.1 points
WMI34.122About 4.95About 19.4 points

Only the published SEM column is an ACIS figure. The projected columns are the 1910 formula applied to it, under the assumption that the removed subtests were as good as the retained ones. They are shown because the direction is not in doubt even if the decimals are: a two indicator index is meaningfully less precise than a five indicator index, and a Quick administration should be read as an estimate with a wider band rather than as a Full Scale result obtained faster. The full 20 subtest case is on the Full Scale page and the comprehensive battery page.

11 Choosing Between Short and Long Without Fooling Yourself

The choice is not between a good test and a bad one. It is between a narrow estimate you can have today and a profile that costs more time, and the right answer depends entirely on what the number has to do afterward.

What you actually wantShortest form that answers itWhy
A rough sense of where you sit overallA short normed form of 15 to 45 minutesGeneral ability is the one thing a short instrument estimates well, since g dominates every composite
To know whether your verbal ability outruns your spatial abilityA full battery with three or more indicators per domainSpecific factors carry too little reliable variance at two indicators to be interpreted, as the WASI-II reanalysis shows
To know whether you clear a stated thresholdA full battery, or a proctored administration if an institution is involvedA wide interval makes a threshold question unanswerable regardless of what the point estimate says
A score somebody else will acceptNone of the above. A proctored session with a licensed psychologistInstitutional acceptance rests on controlled conditions and verified identity, not on test length
To see what the item formats are like before committingA free trial or a free quizFormat familiarity is a legitimate goal and needs no norms at all

The row that trips people up is the third one. A person deciding whether they sit above a line has the strongest possible reason to want an answer quickly and the strongest possible reason not to accept a quick one. Every point of standard error is a point of ambiguity in exactly the region they care about, and short tests are widest precisely where thresholds are set, out at the ends of the scale where the norm sample thins.

The row underneath it is the one people wish were different. No unsupervised online instrument, at any length or price, produces a result an institution will act on. That is a statement about the administration conditions rather than about the psychometrics, and it applies to this site as much as to any other. The routes that do work are compared on the professional against online page, and the named instruments an examiner would use are on the page on what a Stanford-Binet administration actually requires.

For everyone else, the honest framing of a short paid form is the one the technical manual already uses. It is an efficient estimate of selected core abilities. It is not a complete cognitive profile. Those two sentences should appear on the product, not only in the manual, which is why they appear here. What a longer session involves is described on the adult test page.

12 Why Two Short Tests Do Not Add Up to One Long One

The obvious workaround is to take several short tests and average them, and it does not work, for reasons a test publisher documented in detail. Xiaobin Zhou and Susan Engi Raiford wrote Pearson Technical Report Number 2 in November 2011 to address exactly this situation: what happens when a person takes a brief measure and then a full battery containing similar subtests.

They name four contaminating mechanisms. Procedural learning, meaning the strategy knowledge a person picks up from a similar earlier task. Variation in effort, since boredom or discouragement follows a repeat of something already done. Regression to the mean, since extreme first observations move toward the middle on a second measurement. And the Flynn effect, since older norms inflate scores. The first of those is the one they show can be controlled, and their solution is to substitute the brief subtest scores directly rather than re-administer similar tasks.

Their samples are small enough to name honestly. Ninety two people took the WASI-II first and the WAIS-IV second, with a mean interval of 24 days and a range of 13 to 73 days. A second group of 90 took them in the opposite order, with a mean interval of 21 days. Comparing the second group against matched controls drawn from the WAIS-IV normative sample, Full Scale IQ rose by 0.6 points, verbal comprehension by 0.7, and perceptual reasoning by 0.5. None of those increases was statistically significant or of notable effect size, which let the authors conclude that Flynn effect inflation was small in this window and isolate procedural learning as the mechanism to manage.

The practical reading for a consumer is unglamorous. Two short tests taken a week apart are not two independent measurements. The second one is contaminated by the first in ways that push the number up, and the direction of the bias is the flattering one. Averaging two such scores does not narrow the interval by the square root of two, because the errors are correlated rather than independent. The only clean route to a narrower interval is more distinct measurement in a single administration, which is what a long battery is.

A limit worth stating about everything aboveThe ACIS figures quoted throughout this page describe its documented adult reference frame. The assessment is self-administered. Reliability describes consistency inside the frame it was computed in. ACIS is not a clinical instrument, it is not diagnostic, and it is not appropriate for hiring, accommodations, or society admission at any tier. What those boundary conditions mean is set out on the reliability and validity page.

13 Sources Behind This Page

Every coefficient above comes from a document a reader can open, and every projected figure is labeled as a projection in the sentence that uses it. Prices and product listings were read in September 2026 and are dated where they appear.

  • Spearman, C. (1910). Correlation calculated with faulty data. British Journal of Psychology, 3, 271 to 295. The prediction formula relating test length to reliability, published in the same volume as Brown's independent derivation.
  • Brown, W. (1910). Some experimental results in the correlation of mental abilities. British Journal of Psychology, 3, 296 to 322. The second independent derivation of the same relationship.
  • Smith, G. T., McCarthy, D. M., and Anderson, K. G. (2000). On the sins of short-form development. Psychological Assessment, 12(1), 102 to 111. Two general and nine specific methodological failures in short form construction, and the argument that validity does not transfer from a parent form by inheritance.
  • Axelrod, B. N. (2002). Validity of the Wechsler Abbreviated Scale of Intelligence and other very short forms of estimating intellectual functioning. Assessment, 9(1), 17 to 23. A mixed clinical sample of 72 participants, WASI FSIQ-4 correlating .92 and FSIQ-2 correlating .87 with WAIS-III Full Scale IQ.
  • Pierson, E. E., Kilmer, L. M., Rothlisberg, B. A., and McIntosh, D. E. (2012). Use of Brief Intelligence Tests in the Identification of Giftedness. Journal of Psychoeducational Assessment, 30(1), 10 to 24. WASI standardization at 2,245, the two and four subtest reliabilities, and the argument that measurement error can mislabel a person in either direction near a cutoff.
  • McGeehan, B., Ndip, N., and McGill, R. J. (2017). Exploring the Multidimensional Structure of the WASI-II. Archives of Assessment Psychology, 7(1), 7 to 27. Adult sample of 1,200, general factor at 61.2 percent of total variance, omega hierarchical of .84 for g against .16 for each group factor.
  • Crawford, J. R., Parker, D. M., Stewart, L. E., Besson, J. A. O., and De Lacey, G. (1989). Prediction of WAIS IQ with the National Adult Reading Test: Cross-validation and extension. British Journal of Clinical Psychology, 28, 267 to 273. Cross validation sample of 151 aged 16 to 88, 66 percent of Full Scale IQ variance predicted, and a standard error of estimate of 7.68 points on the combined sample of 271.
  • Ammons, R. B., and Ammons, C. H. (1962). The Quick Test (QT): Provisional manual. Psychological Reports, 11, 111 to 162. The original publication of a picture vocabulary estimate of ability. This article could not open the full manual, so no coefficient from it is quoted above.
  • Skues, J. L., and Cunningham, E. G. (2013). An alternative approach to identifying students with learning disabilities in Australian schools. Australian Journal of Educational and Developmental Psychology, 13, 16 to 27. KBIT-2 structure, item counts, and administration time as reported from Kaufman and Kaufman, 2004.
  • Zhou, X., and Raiford, S. E. (2011). Using the WASI-II with the WAIS-IV, Technical Report Number 2. Pearson. The four contamination mechanisms in sequential testing, the linking sample sizes and intervals, and the matched control comparison.
  • Pearson Assessments. WASI-II product listing and KBIT-2 product listing. Administration times, age ranges, qualification levels, prices, and the publishers' own framing of each instrument as a screener.
  • ACIS. Technical manual. All nine published composite reliabilities and standard errors, the subtest g loadings, the index intercorrelations, and the form matrix defining which scores each of the three tiers can produce.

The professional framework behind all of it is explicit. The Standards for Educational and Psychological Testing, published jointly in 2014 by the American Educational Research Association, the American Psychological Association, and the National Council on Measurement in Education, require that measurement error be reported alongside any score or classification, that a score interpretation be supported by evidence for the specific proposed use, and that the composition and limits of the norm sample be disclosed. The APA standards on test use place the responsibility for appropriate interpretation on the person doing the interpreting, a point the WASI-II reanalysis quotes directly from the Standards at page 141.

Read against that framework, the question this page opened with resolves cleanly. A short test is defensible when it publishes what it gives up. The WASI-II does. The NART does. The KBIT-2 does. A five minute page with no reference sample and a certificate at the end does not, and its brevity is the least of the reasons why.

14 Frequently Asked Questions

Is a quick IQ test accurate?

It depends entirely on whether it was normed as a short test. A well built 15 minute form like the WASI FSIQ-2 correlated .87 with the WAIS-III Full Scale IQ in Axelrod's 2002 clinical sample. A free five minute quiz with no reference sample has no accuracy figure at all, because accuracy is measured against something.

What is the shortest IQ test that still means something?

Among documented instruments, the two subtest WASI form at about 15 minutes and the 50 word National Adult Reading Test at two to three minutes both have published relationships with full batteries. The NART's standard error of estimate is 7.68 points, which is useful clinically and far too wide for an individual claim.

How does test length affect reliability?

Through the Spearman-Brown formula, published independently by Spearman and Brown in the British Journal of Psychology in 1910. Multiply length by k and the new reliability equals k times r divided by one plus k minus one times r. Halving a long test costs little. Cutting it to a fifth costs a great deal.

What is a standard error of measurement?

It is the spread of the band around your reported score, and on the IQ metric it equals 15 times the square root of one minus the reliability. A standard error of 2.40 gives a 95 percent interval about 9 points wide. A standard error of 8.2 gives one about 32 points wide.

Can a 5 minute test measure intelligence?

It can produce a rough estimate of general ability if it has a reference sample, and it cannot produce a profile. Reasoning items take about a minute each, so five minutes holds roughly five items, which is too few to span a difficulty range wide enough to place anyone reliably.

Why do short tests measure general ability better than specific abilities?

Because after the general factor takes its share of the variance, two indicators leave almost nothing reliable behind. McGeehan, Ndip, and McGill found omega hierarchical of .84 for g on the WASI-II adult sample against .16 for each of its two group factors.

What is the difference between a short form and a shortened test?

A short form is designed and standardized as a short instrument with its own norms and error terms. A shortened test is a long battery with items deleted, scored against the long test's tables. Smith, McCarthy, and Anderson called the assumption that validity transfers between the two a documented failure of the literature.

Does removing items change more than the length?

Yes. Deleting alternate items widens the difficulty gap between adjacent items, changes the order in which tasks are met, and changes the content sampled. Those are changes in what the test measures, and the Spearman-Brown formula does not model any of them.

How long does the WASI-II take?

Pearson lists 30 minutes for the four subtest form and 15 minutes for the two subtest form, for ages 6 through 90. The kit is listed at 530.40 dollars at qualification Level C, which means it is not sold to the general public.

How reliable is the WASI-II?

Its manual reports internal consistency from .87 to .92 across the four subtests, .94 to .95 for the index scores, and .96 to .97 for the Full Scale IQ, on a norming sample of 2,300 people aged 6 to 90 split as 1,100 children and 1,200 adults.

What is the KBIT-2?

A brief individually administered test of verbal and nonverbal intelligence for ages 4 through 90, published by Kaufman and Kaufman in 2004, taking about 20 minutes. It has three subtests: Verbal Knowledge with 60 items, Riddles with 48, and Matrices with 46.

What is the National Adult Reading Test?

Fifty short words of irregular pronunciation, read aloud and scored for errors, developed by Nelson in 1982. It is used to estimate premorbid ability because reading of irregular words resists many neurological and psychiatric conditions that lower other scores.

How well does the NART predict a full IQ?

Crawford and colleagues reported in 1989 that Nelson's equations predicted 66 percent of the variance in WAIS Full Scale IQ in a cross validation sample of 151 adults. On a combined sample of 271, the standard error of estimate for predicted Full Scale IQ was 7.68 points.

What is the Ammons Quick Test?

A passive response picture vocabulary test published by Robert and Carol Ammons in Psychological Reports in 1962, in which the examinee points at one of several drawings. Its validity coefficients across the follow up literature span an unusually wide band, which is itself informative about single task estimates.

Can I take two short tests and average them?

Averaging does not work as expected, because the second administration is contaminated by the first. Pearson's own technical report names procedural learning, variation in effort, regression to the mean, and the Flynn effect as the mechanisms, and the resulting bias runs upward.

What does the ACIS Quick form include?

Six subtests in about 45 minutes for 15 dollars: Similarities and Vocabulary for verbal comprehension, Matrix Reasoning and Figure Weights for fluid reasoning, and Digit Span and Alphanumeric Sequencing for working memory. It covers three of the six domains and produces no Full Scale IQ.

How does the ACIS technical manual describe the Quick form?

As a focused estimate of selected verbal, fluid, and working memory ability that can support a useful screening level profile, and explicitly not the same thing as a 20 subtest Full Scale administration. The manual also states that any report must name which of the three forms was completed.

How much precision does the Quick form give up?

It supplies two indicators per index where the Full Scale form supplies five for verbal comprehension and fluid reasoning and three for working memory. Applying the 1910 formula to the published standard errors projects a verbal comprehension standard error near 3.72 points instead of 2.40, which is a projection rather than a published figure.

Should I take a short test if I am near a threshold?

No. A threshold question is exactly the case a wide interval cannot answer. At a standard error of 2.40 a reported 132 clears 130 with roughly 80 percent probability, and at 8.2 that falls to about 60 percent, which is close to indifference.

Do short tests overestimate or underestimate?

Either, and the asymmetry is in how people respond. An underestimate produces annoyance and a retake. An overestimate produces a belief. Classical test theory also says the best estimate of a true score regresses toward the mean in proportion to reliability, so low reliability moves the center of the band down, not only its edges.

Is any unsupervised online test acceptable to an institution?

No, including this one. Institutional acceptance rests on controlled conditions, verified identity, and a qualified examiner observing the session, none of which is a function of test length or price. ACIS is not a clinical instrument and is not appropriate for diagnosis, hiring, accommodations, or society admission.

Take the assessment

You get a profile, not a number

ACIS measures six CHC domains across 20 subtests and reports each one with its own normed score and confidence interval, so you can see where you are strong and where you are not.

Free trial, no card required. Full report from $15.