Administration Times

How long an IQ test takes, instrument by instrument

Published administration times run from about 20 minutes for a single matrix task to nearly three hours for a full twenty subtest battery. The spread is not marketing. It is arithmetic: fewer items means a wider error band, and this page shows the calculation rather than asserting the conclusion.

A round wall clock with Roman numerals reading just past twelve, in a blurred hospital corridor.
Every duration on this page comes from a publisher document, not from a blog estimate.

0 The Short Answer

A supervised professional IQ test takes roughly 45 to 90 minutes of testing time, a full unsupervised battery takes about two to three hours, and the three minute quizzes advertised in search results take three minutes and produce a number nobody should act on. Those are publisher figures, not estimates. Pearson states 45 minutes for the seven subtest Full Scale IQ on the WAIS-5. Stoelting, an authorized distributor of the Stanford-Binet Fifth Edition, states 50 minutes for its ten subtest battery. The ACIS Full Scale tier, at twenty subtests, averages about 175 minutes.

The honest complication is that a stated administration time is not the length of your appointment. Publishers report examiner facing testing time: the minutes from the first item to the last. A supervised adult session adds intake, rapport, standardized instruction reading, hand scoring of open ended responses, breaks and behavioral observation, so the door to door figure is routinely double the published number. Unsupervised delivery removes most of that overhead and adds none of its own, which is why an online battery can be longer in testing minutes while taking less of your day.

The second complication is that duration is not a dial the publisher turns freely. It follows from item count, and item count sets the width of the confidence interval around your score. That relationship is a formula, not an opinion, and the second half of this page works through it with real numbers. This page owns total duration. The separate question of why individual items carry time limits at all belongs to its own page and is not repeated here.

45 minutes

Pearson's stated time for the seven subtest WAIS-5 Full Scale IQ.

50 minutes

Stoelting's stated time for the ten subtest Stanford-Binet 5 battery.

175 minutes

Average completion time for the twenty subtest ACIS Full Scale tier.

1 What the Publishers Actually State

Almost every article answering this question repeats numbers it found in another article, so the figures drift. The table below carries only durations that appear on a publisher or authorized distributor page, with the source named in the row. Where a publisher does not state a time, the row says so instead of filling the gap with a guess.

InstrumentWhat the time buysStated administration timeSource of the figure
WAIS-5 (Pearson, 2024)Full Scale IQ from 7 subtests45 minutesPearson product page
WAIS-5 (Pearson, 2024)All 5 primary index scores, 10 subtestsAbout 60 minutesPearson product page
WAIS-IV (Pearson, 2008)Core subtests70 minutesPearson comparison flyer
WAIS-IV (Pearson, 2008)All subtestsAbout 90 minutesPearson comparison flyer
Stanford-Binet 5 (Roid, 2003)One subtest5 minutes, typicalRiverside Insights product page
Stanford-Binet 5 (Roid, 2003)Full Scale IQ from 10 subtests50 minutesStoelting product page
Raven's Standard Progressive Matrices and SPM Plus60 items in five sets, untimed20 to 45 minutesPearson Clinical UK product page
Raven's 2 Clinical Edition (Pearson, 2018)Short form or long form, digital or paperNot stated on the product pagesChecked on the US, Canadian, UK and Asian stores
ACIS Quick6 subtests across 3 domainsAbout 45 minutesACIS published estimate
ACIS Optimized13 subtests across 5 domainsAbout 110 minutesACIS published estimate
ACIS Full Scale20 subtests across 6 domains, Full Scale IQAbout 175 minutesACIS published estimate

One reason these numbers drift in secondhand accounts is that publishers are quoting different units. A figure attached to an instrument may describe a single form, a single composite, one age band, or the entire kit, and none of those is interchangeable with the others. The Raven's matrices are quoted for a 60 item paper form in the United Kingdom and sold as a digital item bank in the United States. The Stanford-Binet 5 spans ages 2 to 85 and older, and a preschooler does not sit the same number of items as a sixty year old. Editions get conflated as well: the 70 minute figure still circulating for the Wechsler adult scale belongs to the fourth edition of 2008, not to the current one. Before reusing any duration you read somewhere, establish which form, which composite, which age band and which edition it was attached to. The last two of those are exactly the fields the inventory that carries one row per instrument pins down, printing the publisher, the edition year and the age range in years and months for every battery named above, which is the minimum a duration needs before it can be attached to anything.

Three patterns fall out of that table. First, the individually administered clinical batteries cluster between 45 and 90 minutes, because they are paced by an examiner who has other things to do that day. Second, a single matrix task sits near 20 to 45 minutes whether it is called Raven's or something else, because one homogeneous task reaches its useful precision quickly and then stops improving. Third, the unsupervised batteries are longer in raw minutes, which is the trade a candidate makes for not having an examiner across the table.

A verified absencePearson publishes the age range and the qualification level for the Raven's 2 Clinical Edition on its store pages but does not publish a completion time on any of the four regional stores checked for this article. A purchaser who needs that figure has to take it from the manual. It is not a number this page will invent, and it is not one you should trust from a blog.

Context for the money side of the same decision sits on the page covering what each of these instruments costs, and the structural comparisons live on the pages for the WAIS-5 and its index structure and the routing design of the Stanford-Binet 5.

2 The WAIS-5 and Where Its Minutes Went

Pearson states two different times for the WAIS-5 because the WAIS-5 is not one test, it is a menu. The Pearson product page gives 45 minutes for the seven subtest Full Scale IQ and about 60 minutes for the ten primary index subtests, and it says plainly that administration time is determined by the composite scores the practitioner wants. Ask for a Full Scale IQ only, and you are in the 45 minute case. Ask for five index scores as well, and you are in the 60 minute case. Add secondary subtests for a clinical question and you go past both.

Compare that with the edition it replaced. Pearson's own WAIS-IV versus WAIS-5 comparison flyer lists 70 minutes for the WAIS-IV core subtests and about 90 minutes for all of them. The Full Scale IQ moved from ten subtests to seven, and the flyer states that the FSIQ can now be obtained about 20 minutes faster. That is where the time went: not into faster items, but into a shorter path to the composite. The full instrument did not shrink. The flyer lists ten primary and ten secondary subtests for the WAIS-5, which is more separately scored tasks than the WAIS-IV offered, alongside five primary index scores and fifteen ancillary ones.

This is the single most important thing to understand about published durations. A shorter stated time can mean a smaller instrument, or it can mean the same instrument with a cheaper route to the headline number. Those are opposite situations for a test taker. On the WAIS-5 it is the second: the battery got broader while the advertised path got shorter, and the practitioner decides which one you actually sit. The specifics of that change are covered on the page comparing the fourth and fifth editions.

One detail from the same Pearson page deserves its own line, because it explains why two people booked for the same 45 minute test can be in the room for wildly different lengths of time. Pearson notes that examinees who are intellectually gifted commonly require long testing times, because they meet discontinue rules late in the item order or not at all, and that the new later start points for suspected giftedness cut testing time by roughly 25 percent relative to standard start points. Duration, in other words, is partly a property of the person being tested. That mechanism gets a full section below, and the composite it feeds is described on the page on what a Full Scale IQ requires.

3 The Stanford-Binet 5 and the Raven's Matrices

The Stanford-Binet 5 publishes its time per subtest rather than per battery, which is a more honest way to describe a routed instrument. Riverside Insights, the publisher, states on its product page that typical administration time is five minutes per subtest for examinees aged 2 to 85 and older. Stoelting, an authorized distributor, states the battery figure explicitly: 50 minutes, described as approximately five minutes for each of the ten subtests, five nonverbal and five verbal.

Per subtest reporting matters here because the Stanford-Binet 5 routes. Two routing subtests establish an approximate ability level, and the remaining subtests then start at an appropriate difficulty point rather than at item one. The Abbreviated Battery IQ is built from those two routing subtests alone. Neither publisher page states a separate time for the abbreviated form, so this page will not supply one, but the structural consequence is clear enough: the abbreviated battery buys a large reduction in minutes by measuring two tasks instead of ten, and it pays for that in the width of the interval around the resulting score. That trade is quantified further down this page.

The Raven's family sits at the other end of the design space. Pearson Clinical UK states the completion time for the Standard Progressive Matrices and SPM Plus as untimed, individually or in groups, 20 to 45 minutes, across 60 items arranged in five sets of twelve. Raven's 2, the current clinical edition authored by Raven, Rust, Chan and Zhou and published in 2018, exists in a paper form, a digital long form and a digital short form, and Pearson's public store pages give the age range of 4 to 90 and the qualification level without stating minutes for any of the three.

Why a matrix test is shortA single well constructed matrix task reaches useful precision quickly because every item measures nearly the same thing. That homogeneity is exactly what makes it fast and exactly what limits it: the score describes one slice of ability rather than a profile. McLeod and McCrimmon reviewed Raven's 2 in the Journal of Psychoeducational Assessment in 2021, volume 39 issue 3, and characterized it as a quick and efficient estimate of general cognitive ability, with the qualifications that phrase implies. The record is indexed in ERIC.

Readers deciding between a Wechsler style battery and a Binet style routed instrument will find that comparison on the page setting the two side by side, and the mechanics of the item type itself on the matrix reasoning subtest page.

4 The Three ACIS Durations and What Buys the Difference

ACIS publishes three completion estimates because it sells three different amounts of measurement, and the minutes map onto subtests almost linearly. Quick is six subtests across three of the six domains and averages about 45 minutes. Optimized is thirteen subtests across five domains and averages about 110 minutes. Full Scale is all twenty subtests across all six domains and averages about 175 minutes, roughly two hours and fifty five minutes. Those are average completion estimates based on the current subtest mix and stopping rules, so individual pace varies, and the sections below explain why the variance is large rather than small.

TierSubtestsDomains coveredAverage completionMinutes per subtestProduces a Full Scale IQ
Quick6 of 203 of 6 (VCI, FRI, WMI)About 45 minutesAbout 7.5No
Optimized13 of 205 of 6 (all but QRI)About 110 minutesAbout 8.5No
Full Scale20 of 206 of 6About 175 minutesAbout 8.8Yes

The near constant minutes per subtest column is the point. Nothing is being padded and nothing is being rushed. Each subtest costs what it costs, and the tiers differ by how many of them you sit. The last column is the one that decides the purchase for most readers, because a Full Scale IQ is a composite over all six domains and cannot be produced from a partial battery. Buying 45 minutes buys a real profile across three domains. It does not buy a Full Scale number, and no honest instrument will hand you one from six subtests.

Two practical facts change how those minutes feel. Progress is saved between subtests, so the battery does not have to be finished in one sitting, and the assessment stays open for thirty days. That makes 175 minutes a very different proposition from 175 minutes in a clinic chair. It also introduces a variable a supervised session does not have, since a battery spread across four evenings is administered under four different sets of conditions. The domains those minutes are spread across are described on the cognitive domains page, and the theoretical structure behind them on the CHC model page.

ACIS is an unsupervised online assessment, not a clinical instrument, and it is not appropriate for diagnosis, hiring decisions, accommodation requests or high IQ society admission. Its adult reference frame is a modelled frame built from 3,243 English speaking records aged 16 to 90 rather than a census sample, and the people in it selected themselves by choosing to take a test. Reading the resulting numbers is covered on the score interpretation page.

5 Why Item Count Sets the Error Bar

Test length is not a comfort setting. It is the input to the formula that decides how wide your confidence interval is. The standard error of measurement is the standard deviation of the score scale multiplied by the square root of one minus the reliability coefficient. On the IQ metric the standard deviation is 15, so the whole calculation is SEM equals 15 times the square root of one minus r. Nothing else enters it. Every additional good item raises r, and every rise in r narrows the band around the number you are given.

Reliability of the compositeStandard error of measurement95 percent intervalWidth of that interval
.991.5 pointsPlus or minus 2.95.9 points
.953.4 pointsPlus or minus 6.613.1 points
.904.7 pointsPlus or minus 9.318.6 points
.806.7 pointsPlus or minus 13.126.3 points
.708.2 pointsPlus or minus 16.132.2 points
.609.5 pointsPlus or minus 18.637.2 points

Run the ACIS composite through it as a check. The published Full Scale IQ composite omega is .9886. One minus that is .0114, whose square root is .1068, and fifteen times .1068 is 1.60. That reproduces the published ACIS standard error of measurement of about 1.60 IQ points exactly, which is the point of showing the arithmetic: these are not numbers a test publisher chooses, they are numbers that fall out of how much measurement was done. At that standard error, a 95 percent interval spans about 6.3 points, so a reported 128 sits in a band of roughly 125 to 131.

Now read the bottom half of the table. At a reliability of .70, which is a perfectly ordinary figure for a short homogeneous quiz, the 95 percent interval is 32 points wide. A person reported at 120 cannot be distinguished from a person reported at 104 or 136. The number still appears on the screen in the same font. What has changed is that it no longer separates anyone from anyone.

The usual objection is that reliability and length are separate things, and in principle a short test of superb items could beat a long test of poor ones. In practice item quality has a ceiling and length does not, which is why every serious battery is long. The ACIS reliability figures come with their own boundary conditions, since the analysis set is 2,750 complete records from a self selected sample administered without supervision. Those conditions are laid out on the page on measurement quality, and the metric itself on the page explaining the fifteen point standard deviation.

6 The Arithmetic of Making a Test Shorter

There is a formula that predicts what happens to reliability when you cut a test's length, and applying it to a real battery is the fastest way to see why durations are what they are. The Spearman-Brown prediction formula, published independently by Spearman and by Brown in 1910, says that if you multiply a test's length by a factor k, the new reliability equals k times r divided by one plus k minus one times r. It assumes the items you keep are as good as the items you cut, which makes every result below a best case rather than a forecast.

Take a battery with a composite reliability of .98 that runs 175 minutes, and shorten it. The table applies the formula and then converts each result into an error band on the IQ metric using the standard error calculation from the previous section.

Fraction of the battery keptApproximate minutesPredicted reliabilityStandard errorWidth of the 95 percent interval
All of it175.9802.18.3 points
One half88.9613.011.6 points
One quarter44.9254.116.2 points
One tenth18.8316.224.2 points
One twentieth9.7108.131.7 points
One sixtieth3.44611.243.8 points

Read down the reliability column and notice how gently it falls at first and how fast it collapses at the end. Halving a long battery costs very little, which is exactly why abbreviated forms of good instruments are defensible for screening. Cutting to a tenth costs a great deal. Cutting to a sixtieth destroys the measurement: at .45 the score is closer to a coin flip than to an assessment, and the 95 percent interval is wider than the distance from the fifth percentile to the ninety fifth.

This is not a fringe argument. Kruyen, Emons and Sijtsma reviewed the shortened test literature in the International Journal of Testing in 2013, volume 13 issue 3, and found that shortening a test can substantially affect both the reliability and the validity of the resulting scores, and that test constructors and users frequently fail to address those consequences. Their record is public at the Tilburg University research portal. The practical reading is on the page on what makes an online test accurate and the comparison of free tests against validated ones.

7 Why Three Minutes Cannot Produce a Usable Score

Work it from the other direction and the conclusion is the same, which is the useful thing about arithmetic. Give a three minute test the most generous assumptions available. Suppose the items are so efficient that a person answers one every ten seconds, giving eighteen items. Suppose every item is a well written measure of the same ability, with no guessing, no misread instruction and no distraction. Standardized alpha for a set of homogeneous items is k times the mean inter item correlation, divided by one plus k minus one times that correlation, where k is the number of items.

At a mean inter item correlation of .15, which is respectable for ability items, eighteen items yield an alpha of about .76, a standard error of about 7.3 points, and a 95 percent interval about 29 points wide. Push the item quality to an unusually high mean correlation of .25 and you reach about .86, a standard error of 5.7, and a band about 22 points wide. Drop it to .10 and you fall to about .67, a standard error of 8.7, and a band about 34 points wide. The Spearman-Brown route in the previous section landed at 44 points. Every path converges on the same territory: a three minute test carries an uncertainty of roughly twenty to forty IQ points.

An interval that wide does not distinguish average from superior. It does not distinguish superior from average either, which is the direction that actually hurts people, because the flattering result is the one that gets believed. Measured against the alternative, those three minutes buy less than they appear to: a bare self estimate correlates about .30 with a measured score and carries an interval near 55 points, so eighteen items close roughly half the distance between guessing and measuring and leave the rest of it open, which is the comparison worked through on what a self estimate is worth in points. And this is still the optimistic case, since it assumes the test has a norm sample at all. Most three minute tests do not. A raw score becomes an IQ only by comparison with a documented reference group of known composition and known age structure, which is a research undertaking rather than a scripting one, as the page on norming sets out.

What a three minute test is good forIt is a legitimate curiosity object and a legitimate marketing device, and there is nothing wrong with taking one. What it cannot do is support a decision, a self concept or a claim to anybody else. If a site reports a three minute result to two decimal places, the precision is decorative. The relevant reading is on the page on online test accuracy and the page on free tests and their limits.

The uncomfortable corollary applies to serious instruments too. Nobody should treat a single reported number as exact, including a number from a long battery. The difference is the size of the band: about six points on a full ACIS Full Scale IQ, about thirty to forty on a three minute quiz. Converting a score to a rank is covered on the percentile calculator page, and how uncommon a given score is on the rarity page.

8 Why Two People Take Very Different Amounts of Time

The same test, administered correctly to two people on the same day, can take twice as long for one as for the other, and the reason is built into the instrument on purpose. Ability tests order items by difficulty and stop presenting them once a person has clearly passed their ceiling. That mechanism is called a discontinue rule. It exists to spare people the demoralizing experience of failing item after item, and to spare everyone the wasted minutes of measuring something already established. Its side effect is that duration becomes a function of the person, not just of the test.

Pearson states the consequence directly for the WAIS-5: examinees who are intellectually gifted commonly require long testing times, because they meet discontinue rules late in the item order or not at all. That is why the fifth edition added later start points for suspected giftedness, which Pearson reports cut testing time by about 25 percent for those examinees. A high ability adult sitting a nominally 45 minute battery is not being mistreated by a slow examiner. They are simply reaching more items.

The published ACIS administration specification shows the same logic in an unsupervised setting, with different implementations across the twenty subtests.

SubtestPublished administration ruleEffect on duration
Vocabulary45 words, 1:30 per itemCeiling of 67.5 minutes if every item is reached and every second used
Similarities31 pairs, 1:30 per itemCeiling of 46.5 minutes on the same arithmetic
Arithmetic26 audio problems, 30 seconds each, ends after 3 errorsHighly variable, often a fraction of the ceiling
Digit SpanThree parts, each ends after 2 consecutive errorsScales with span, so it is short for most people
Spatial Navigation30 map items, 30 minutes total, no discontinue ruleEffectively fixed length for everyone
Coding120 seconds, up to 135 symbolsIdentical for everyone by construction

Multiply out the top two rows and something striking appears. Vocabulary and Similarities alone have a combined ceiling of 114 minutes, yet the whole twenty subtest battery averages about 175. Almost nobody approaches the ceilings, because most people answer a vocabulary item in a fraction of ninety seconds and because discontinue rules end several subtests early. The published average is an average over people, and the distribution around it is wide. A short administration is not a broken one, and a long administration is not evidence of struggle. Both are the rule working. The individual tasks are described on the subtest types page, with the fixed length navigation task on its own page and the graded verbal task on the vocabulary page. Readers whose reason for asking is a suspected high score will find the relevant framing on the gifted range page.

9 Speed Tasks and Power Tasks Cost Different Amounts of Time

Two subtests can measure with equal precision while one takes two minutes and the other takes forty, because they are answering structurally different questions. A speed task asks how many easy items a person completes in a fixed window. The window is the measurement. A power task asks how difficult an item a person can solve given room to think. There the measurement is the difficulty reached, and reaching it requires climbing a ladder of items, which takes minutes that cannot be compressed.

The ACIS Processing Speed Index illustrates the first case at its cleanest. Coding runs 120 seconds with up to 135 symbols available, and Symbol Search runs 120 seconds across 80 trials. The entire index is four minutes of timed work inside a battery of nearly three hours. Extending either task to ten minutes would not sharpen it. It would import fatigue and motivation into a measure of clerical rate, changing what the subtest measures rather than how well it measures. Short here is correct, not cheap.

The power tasks in the same battery run the other way. Vocabulary allows ninety seconds per item across 45 words. Mathematical Achievement allows 35 minutes for 25 problems. Logic Grid allows 20 minutes across 26 grids with no visible countdown. Each of those needs many items spanning a wide difficulty range, because precision at the top of the scale requires items that only people at the top of the scale can solve, and items like that cannot be answered in eight seconds by anybody.

The boundary of this pageThis section covers only the effect of the speed and power distinction on total duration. Why tests impose limits in the first place, how per item limits are chosen, what time bonuses do, and whether speed and power should be separated as constructs are all handled on the dedicated page on time limits in cognitive testing. Nothing there is repeated here, and nothing here is repeated there.

The design consequence for anyone comparing two products is worth stating plainly. If a battery reports a processing speed score, the minutes it spends on it will be tiny, and that is a sign of competence rather than corner cutting. If a battery reports a reasoning composite in the same tiny window, something is wrong, because there is no way to place a person on a difficulty ladder without presenting enough rungs. Duration by itself proves nothing. Duration read against what the test claims to measure is highly diagnostic. The construct is described further on the processing speed page, with task specifics on the coding page and the symbol search page.

10 What Supervision Adds to the Clock

The gap between a publisher's stated administration time and the length of a real appointment is not padding, it is the work that makes a supervised score defensible. Pearson's 45 minutes for a WAIS-5 Full Scale IQ describes the interval from the first item to the last. Around it sits an entire session that the figure does not include, and understanding what is in there explains most of the confusion people have when they book an assessment and are told to set aside a morning.

A supervised session begins with intake and history taking, because a score without context is not interpretable and a clinician needs to know what they are looking at. It continues with rapport building, which is not a courtesy but a measurement condition, since a person who is anxious or guarded produces a score that underestimates them. Instructions are read verbatim from the manual, because standardization is what makes the norms apply. Teaching items and sample items are administered and, where the manual allows, corrected. Manipulatives are set out and cleared away. Open ended verbal responses are queried when ambiguous and scored by hand against criteria that require judgment. Breaks are offered. Behavior is observed and recorded throughout, because a valid score depends on the examinee having actually engaged.

Add the report afterwards, which the examinee never sees the clock on, and a 45 minute instrument becomes a session of two to three hours plus several hours of professional time. Those hours are billed by somebody who cleared a Level C purchase gate, normally a doctorate with formal training in assessment or a state license, and that gate is only the first of three independent ones, since scope of practice and the acceptance of the report by whoever receives it are decided separately, as the page on who is permitted to administer one sets out. The full sequence from referral to report is set out on the page describing the supervised process, and the comparison of the two delivery models on the page setting supervised against online administration.

Unsupervised delivery removes almost all of that overhead and cannot replace what it removes. There is no rapport, no querying of an ambiguous answer, no observation of whether the person was engaged, and no examiner to notice that the room was noisy. What it gains is scheduling freedom, identical instructions for everyone by construction, exact timing, and automatic scoring with no transcription error. The International Test Commission's guidelines on computer based and internet delivered testing, available through the Commission's guidelines page and summarized by the British Psychological Society, treat supervision level as a defining property of an assessment rather than a detail, and they are right to. It is why ACIS states its unsupervised status wherever a score is reported, and why it is not offered for diagnosis or selection. Practitioners administering it to participants under their own supervision use the ACIS Professional workspace, described on the section covering supervised administration. Booking options for a formal assessment are listed on the page on where to take one.

11 Where the Minutes Actually Go: Breadth Against Depth

Any battery spends its minutes on exactly two things, and knowing the split tells you what the resulting score can be used for. Depth is precision within a narrow ability: more items on the same construct, a tighter interval on that one number. Breadth is coverage across constructs: more domains sampled, a profile instead of a point. The two compete for the same clock, and every published duration is a decision about how to divide it.

The Raven's matrices spend nearly everything on depth in one place. Sixty items, one item type, 20 to 45 minutes, and a single score that is precise about a narrow band of reasoning and silent about everything else. That is a coherent design, and for many research uses it is the right one. The Wechsler batteries and ACIS spend their minutes the other way. Twenty separately scored tasks across six domains cost a great deal more time and return a profile in which a verbal comprehension score and a processing speed score can differ substantially for the same person.

That difference is not noise, and it is the main thing a longer administration buys. A single number describes level. A profile describes shape, and shape is where the practical information usually lives, because a person whose reasoning score sits well above their processing speed score has a different working life from a person with the reverse pattern, even when the two composites are identical. Producing shape reliably requires enough items in every domain to make each index trustworthy on its own, which multiplies the minutes by the number of domains rather than adding to them.

The ACIS structure shows the cost concretely. Twenty subtests feed six domain indices and a Full Scale composite whose published g loading is .958 against a higher order model fitting at CFI .9761, TLI .9726, RMSEA .0406 and SRMR .0217, estimated on 2,750 complete records. Those figures describe a self selected sample tested without supervision against a modelled adult reference frame, not a census sample, and they should be read with that boundary attached. What they establish is that the composite behaves as a general factor measure, which is the return on the third hour. Background on that factor is on the general factor page, and on the constructs a battery samples on the page on what intelligence tests measure.

The practical rule that falls out of this is short. If you want a level, a shorter well built instrument will serve. If you want a shape you can act on, you are buying domains, and domains are bought in minutes. The delivery format for the longer option is described on the online battery page.

12 Planning the Session You Are Actually Going to Take

Knowing that a battery averages 175 minutes is useless unless you also know how those minutes should be arranged, and for an unsupervised test that arrangement is entirely your responsibility. In a clinic the examiner manages pacing, notices fatigue and offers breaks at defensible points. Alone at a desk, nobody does that for you, so the scheduling decisions a clinician would make quietly become decisions you have to make deliberately.

Fatigue is the first one. Sustained reasoning across three hours degrades in most adults, and the subtests sitting late in a long unbroken session will be measured under conditions the earlier ones were not. Because ACIS saves progress between subtests and keeps the assessment open for thirty days, the sensible pattern is two or three sittings of an hour rather than one heroic afternoon. That protects the later domains from being penalized for arriving last.

Consistency is the second. Splitting a battery across evenings introduces its own variable, since each sitting has its own noise level, alertness and device conditions. The fix is not to avoid splitting but to make the sittings resemble each other: same room, same machine, comparable time of day, and never a subtest started when you are already tired or about to be interrupted. Three subtests in the ACIS battery are delivered by audio, so headphones and a quiet room are a prerequisite rather than a preference before those begin.

The third decision is which tier to start with, and it should follow from the question you are asking rather than from the price. If you want a Full Scale IQ, only the complete battery produces one, and no amount of thinking about it will change that. If you want a defensible read on reasoning and working memory without committing three hours, the shorter tier is honest about being exactly that. Preparation advice that survives scrutiny, which is a short list, is on the preparation page, and the adult specific options on the adult testing page.

One thing not to do is retake a battery immediately because the first duration surprised you. Practice effects are real, the second administration is not independent of the first, and a score obtained that way is harder to interpret than the one you already have.

13 What a Duration Does and Does Not Certify

Length is evidence about a test, and it is weak evidence taken alone. A long test can be long because it has poor items and needs many of them. A short test can be short because it measures one thing efficiently and says so. The claim that duration alone settles quality is exactly as wrong as the claim that it is irrelevant, and the useful position is between the two: duration is interpretable only against what the instrument says it produces.

Three questions do the work. First, what composite is being claimed, and does the item count plausibly support it? A Full Scale IQ from six subtests is a category error regardless of how carefully those six were built. Second, what is the reported standard error, and does the arithmetic in this page reproduce it from the stated reliability? If a publisher reports reliability but no standard error, compute it yourself with fifteen times the square root of one minus r, and see whether the interval they quote matches. Third, against what reference group is the score expressed, and how was that group assembled? A perfectly reliable test with no defensible norm sample still cannot produce an IQ.

Applied to this page's own instrument, that discipline gives a specific answer rather than a promotional one. The ACIS Full Scale tier averages about 175 minutes because it administers twenty subtests across six domains, its published composite omega of .9886 implies a standard error of about 1.60 IQ points, and its reference frame is 3,243 English speaking adult records aged 16 to 90 which are self selected rather than census drawn. Every one of those facts is checkable, and the last one is a limitation stated in the same breath as the first two. The full derivations sit in the published technical documentation.

This is the framework the profession itself uses. The Standards for Educational and Psychological Testing (2014), developed jointly by the American Educational Research Association, the American Psychological Association and the National Council on Measurement in Education, require that test documentation report the conditions of administration, the composition of the norm group, and the precision of the scores, so that a user can judge whether an instrument supports the interpretation being made of it. The APA testing standards page sets out that framework, and it is deliberately silent on how many minutes a test should run. That silence is the answer to the question this page began with. There is no correct duration. There is only a duration that matches the claim, a claim that matches the evidence, and evidence a reader can check, which is the standard any instrument reporting a number about a person should be held to.

14 Frequently Asked Questions

How long does an IQ test take?

It depends entirely on the instrument. Pearson states 45 minutes for the seven subtest WAIS-5 Full Scale IQ, Stoelting states 50 minutes for the ten subtest Stanford-Binet 5, and the twenty subtest ACIS Full Scale tier averages about 175 minutes.

How long is the WAIS-5?

Pearson gives two figures on its product page: 45 minutes if the practitioner wants only the seven subtest Full Scale IQ, and about 60 minutes for the ten subtests that yield all five primary index scores. Secondary subtests add further time.

Why is the WAIS-5 quicker than the WAIS-IV?

Because the route to the composite got shorter, not because the instrument shrank. The Full Scale IQ moved from ten subtests to seven. Pearson's comparison flyer lists 70 minutes for WAIS-IV core subtests against 45 for the WAIS-5 Full Scale IQ.

How long does the Stanford-Binet 5 take?

Riverside Insights states a typical five minutes per subtest. Stoelting, an authorized distributor, gives 50 minutes for the ten subtest full battery. Routing subtests set each examinee's starting difficulty, so nobody works through every item from the beginning.

How long does Raven's 2 take?

Pearson does not state a completion time for Raven's 2 on its US, Canadian, UK or Asian store pages, so this site will not supply one. For comparison, Pearson does state 20 to 45 minutes for the Standard Progressive Matrices.

How long is the ACIS Quick tier?

About 45 minutes on average, covering six subtests across three of the six domains. It returns a genuine profile for those domains. It does not return a Full Scale IQ, which requires the complete twenty subtest battery.

How long is ACIS Optimized?

About 110 minutes on average for thirteen subtests across five domains. It covers everything except the quantitative domain. Like the shorter tier, it produces domain scores rather than a Full Scale composite.

How long is the ACIS Full Scale battery?

About 175 minutes on average, roughly two hours and fifty five minutes, across all twenty subtests and all six domains. That average is over people, and individual completion times vary considerably around it.

Do I have to finish in one sitting?

No. Progress is saved between subtests and the assessment stays open for thirty days. Splitting it into two or three sessions is usually better, because sustained reasoning degrades and the last subtests should not be measured under worse conditions than the first.

Why do two people take different amounts of time?

Discontinue rules stop a subtest once a person has clearly passed their ceiling. Someone who keeps answering correctly reaches more items and therefore sits longer. Pearson notes this explicitly for high ability examinees on the WAIS-5.

Can a three minute test give me a real score?

No. Eighteen items of good quality reach a reliability near .76 at best, which puts a roughly 29 point band around the result. Two independent calculations on this page both land between twenty and forty points of uncertainty.

Does a longer test always mean a better one?

Not by itself. A long test can be long because its items are weak. Duration is interpretable only against the claim being made: a reasoning composite needs many items, while a processing speed measure is correctly finished in two minutes.

How many items does a test need?

Enough for the claim it makes. Precision rises with item count following the Spearman-Brown relationship, steeply at first and then with diminishing returns, so a composite claiming a narrow confidence interval needs a great many items across a wide difficulty range.

What is the standard error of measurement?

The expected spread of a person's scores across repeated administrations. On the IQ metric it equals fifteen multiplied by the square root of one minus the reliability coefficient. Nothing else enters the calculation, which is why it is easy to check.

What interval surrounds an ACIS Full Scale IQ?

The published composite omega of .9886 implies a standard error of about 1.60 IQ points, so a 95 percent interval spans roughly six points. That figure comes from a self selected, unsupervised sample against a modelled adult reference frame.

Is an appointment longer than the stated time?

Almost always. Publisher figures cover first item to last. A real session adds intake, rapport, verbatim instructions, hand scoring, queries on ambiguous answers, breaks and behavioral observation, which commonly doubles the time before any report is written.

Why are processing speed subtests only two minutes?

Because they measure rate, and the window is the measurement. Coding runs 120 seconds and Symbol Search runs 120 seconds. Lengthening either would import fatigue and change what is being measured rather than measure it more precisely.

Do time limits make a test shorter overall?

They cap the worst case rather than shorten the typical case. Most people answer well inside the per item limit, so the caps mainly bound how long an unusually slow administration can run. The construct question behind limits has its own page.

Does fatigue affect scores on a long battery?

It can, which is the strongest argument for splitting an unsupervised session. Subtests taken late in a three hour block are measured under different conditions from those taken first, and that difference shows up as profile scatter rather than as a lower total.

Is the abbreviated Stanford-Binet good enough?

For screening, often yes. The Abbreviated Battery IQ uses the two routing subtests, so it buys a large time saving at the cost of a wider interval around the score. It is not a substitute for the full battery when a decision follows.

Should I just pick the shortest tier?

Only if the shortest tier answers your question. If you want a Full Scale IQ, nothing but the complete battery produces one. If you want a read on reasoning and working memory, the shorter tier is honest about being exactly that.

Take the assessment

You get a profile, not a number

ACIS measures six CHC domains across 20 subtests and reports each one with its own normed score and confidence interval, so you can see where you are strong and where you are not.

Free trial, no card required. Full report from $15.