What is a cognitive test? A standardized sample of mental performance, scored against a reference group
A cognitive test is a standardized procedure that samples how a person reasons, remembers, attends, processes and knows, under fixed conditions, and scores the sample against a reference group. The phrase covers intelligence batteries, college admission tests, hiring screens, neuropsychological evaluations, ten minute dementia screens, research toolkits, school group tests and online batteries. This page defines the term, maps every family, and explains what separates a measurement from a quiz.
Block and card design tasks appear in intelligence batteries, neuropsychological evaluations and dementia screens alike; what differs is the reference group the arrangement is scored against.
0 Quick Answer
A cognitive test is a standardized procedure that samples a person's mental performance, on tasks of reasoning, memory, attention, speed, language or acquired knowledge, under fixed conditions of timing and instruction, and scores that performance against a reference group so that the result is a position rather than a count. Standardized means the items, instructions, time limits and scoring rules are the same for everyone. Sampled means the test observes a few dozen behaviours and infers a broader ability from them, which is why every score carries error. Fixed conditions mean that a stopwatch, a quiet room and a script are part of the instrument. Scored against a reference group means that a raw count becomes meaningful only when it is placed within the distribution of a defined sample: adults of the same age, applicants for a job, or patients of a certain age and education.
The phrase covers several families that share the definition and answer different questions. Intelligence batteries such as the Wechsler Adult Intelligence Scale, Fifth Edition, or the Stanford-Binet 5 estimate general ability with a profile of domain scores. Admission tests such as the SAT, ACT and GRE predict academic performance. Pre-employment tests such as the Wonderlic screen applicants in minutes. Neuropsychological batteries explain a change in functioning after injury or illness. Brief clinical screens such as the Mini-Mental State Examination and the Montreal Cognitive Assessment flag possible impairment in about ten minutes with a total out of 30. Research toolkits such as the NIH Toolbox measure cognition in studies, school group tests such as the CogAT identify students for programs, and online batteries such as ACIS, which we sell, report a profile to the person taking them.
What separates a good test from a quiz is what it publishes: the composition of its norm sample, the reliability of each score, the standard error that follows, and evidence that the score means what the publisher says it means. The Standards for Educational and Psychological Testing, issued jointly by the American Educational Research Association, the American Psychological Association and the National Council on Measurement in Education in 2014, define validity as the degree to which evidence and theory support the interpretation of a score for a proposed use, which is why a test can be sound for one purpose and meaningless for another. The page on what IQ is explains the score most people associate with the phrase; this page maps the whole territory.
30
The maximum total on both the Mini-Mental State Examination of 1975 and the Montreal Cognitive Assessment of 2005, the two brief clinical screens; the MoCA paper proposed 26 as the lower bound of normal.
.31
The validity of general mental ability tests for predicting job performance in the 2022 reanalysis by Sackett and colleagues, revised down from the .51 reported by Schmidt and Hunter in 1998.
16 to 90
The age range in years across which both the WAIS-5 norms and the ACIS adult reference frame of 3,243 records are defined; a cognitive test is only interpretable inside the ages its sample covers.
1.60
The standard error, in IQ points, of the ACIS Full Scale composite at a reliability of .9886, the figure that turns a reported score into an interval.
A cognitive test measures behaviour rather than the brain, and each term in the definition marks a boundary between a measurement and a guess. A sample of performance is what a test observes. A vocabulary subtest does not measure all the words a person knows; it presents thirty or forty of them, chosen to span a range of difficulty, and infers the size and precision of a person's vocabulary from the pattern of answers. A matrix task presents a set of figural patterns and infers fluid reasoning from how many are solved in the time allowed. Because a test is a sample, every score is an estimate with a standard error, and a test that reports a score without one has left out the part that says how far to trust it.
Fixed conditions are what make the sample comparable. The Standards describe standardization as the requirement that all test takers receive the same content, instructions, timing and scoring, or a demonstrably equivalent version of them, so that differences in scores reflect differences in the people rather than in the occasion. A time limit is a design choice with consequences: a speeded task measures how much a person can do per unit of time, while a power task with generous time measures how hard a problem a person can solve at all, and the page on subtest types sorts the common formats along that line. Instructions are scripted for the same reason.
Scoring against a reference group distinguishes a norm referenced test from a criterion referenced one. A driving test is criterion referenced: the standard is fixed and the pass rate of other candidates is irrelevant. Almost every cognitive test is norm referenced: a raw score of 24 correct is neither high nor low until it is placed within the distribution of raw scores in a defined sample, and the sample is therefore part of the instrument. The Standards require publishers to describe that sample, its size, its composition and the year it was collected, because a score refers to it and to nothing else.
The genus and the species matter for search. An IQ test is a cognitive test whose composite is expressed on a scale with a mean of 100 and a standard deviation of 15 against an age based general population sample. A cognitive ability test in the hiring sense, which the page on cognitive ability tests covers, samples the same general factor for a narrower purpose. A dementia screen samples orientation and recall against a cutoff. All three are cognitive tests, only the first produces an IQ, and treating them as interchangeable is the most common error in this subject.
2 The Families and What Each One Answers
Choose a cognitive test by the question it was built to answer, and the family, the examiner, the report and the price follow from the question. The cost column is a band rather than a quote: clinical fees are set by the provider and vary by an order of magnitude, employer and school tests are paid by the organization rather than by the person tested, and the online figures are the ACIS prices at the time of writing. Each band is linked to its source in the section on that family.
Purpose
Family
Example
Who administers it
What it reports
Typical cost band
Estimate general ability with a full profile, for a diagnosis, placement or accommodation
Intelligence battery
WAIS-5, Stanford-Binet 5
Licensed psychologist, one to one
Full Scale IQ, index scores, percentiles and confidence intervals in a signed report
100 to 400 dollars at a university training clinic, several times that privately; the kit costs the professional from 215 dollars
Estimate reasoning with little language
Matrices test
Raven's 2
Psychologist or trained examiner
A standard score and percentile for one ability
Sold to qualified professionals only
Predict academic performance for admission
Achievement and aptitude test
SAT, ACT, GRE
The testing organization, at a centre or under remote proctoring
Section and total scores on the test's own scale, with percentiles among test takers
A registration fee paid by the candidate
Screen applicants for a job
Pre-employment cognitive test
Wonderlic Select, Criteria CCAT
The employer, usually online and unproctored
A job fit report to the employer; no numerical cognitive score to the candidate
Paid per candidate by the employer
Explain a change in functioning after injury, illness or a suspected disorder
Neuropsychological battery
WAIS-5 plus memory, attention, executive, language and validity measures
Clinical neuropsychologist, over several hours
Domain by domain profile, diagnosis or differential, recommendations
Priced in clinician hours: hundreds of dollars at a training clinic, thousands privately
Flag possible impairment in minutes
Brief clinical screen
MMSE, MoCA
Physician, nurse or psychologist during a visit
A total out of 30 against a cutoff; no IQ, no profile
Part of a consultation
Measure cognition consistently across a study
Research toolkit
NIH Toolbox Cognition Battery
Trained research staff on a tablet
Normed scores per test and composites, for the study
Obtained by the institution
Identify students for gifted and other programs
School group test
CogAT
Schools, to whole classes
Verbal, quantitative, nonverbal and composite scores by age and grade
Paid by the district
Know your own profile at a stated precision
Online self administered battery
ACIS
The person, in a browser
Full Scale IQ, six indices, 20 scaled scores, percentiles and standard errors
Five subtests free; 15, 30 or 50 dollars one time
Two axes organize the table. The first is breadth: an intelligence battery or a research toolkit samples six or more domains and reports each, while a screen or a hiring test samples one composite in a few minutes and reports a total. The second is the reader of the result: a clinician, an admissions office, an employer, a study database or the person tested. A test built for one reader is rarely useful to another, which is why a MoCA total says nothing to an employer. The page on types of IQ tests sorts the first family in more detail.
3 Intelligence Batteries: WAIS-5, Stanford-Binet 5, Matrices and ACIS
An intelligence battery is the reference case of a cognitive test, because it samples several domains at once and reports both a general estimate and the profile beneath it. The Wechsler Adult Intelligence Scale, Fifth Edition, is the instrument most adult clinical assessments in the United States are built around. Pearson's product page, read on September 19, 2026, lists a 2024 publication date, an age range of 16 years 0 months to 90 years 11 months, a completion time of 45 minutes for the seven subtest Full Scale IQ and 60 minutes for the ten primary index subtests, a qualification level of C, which restricts purchase to buyers with doctoral training in assessment or an equivalent license, and complete kits from 215 dollars. The page on what the WAIS-5 is walks through its index structure, and the report a psychologist writes from it prints each index with a percentile and a confidence interval.
The Stanford-Binet 5 is the other individually administered battery a psychologist is likely to choose. It organizes ten subtests into five factors, fluid reasoning, knowledge, quantitative reasoning, visual spatial processing and working memory, each tested once verbally and once nonverbally, as the page on how the Stanford-Binet 5 works explains. Both instruments produce composites on a scale with a mean of 100 and a standard deviation of 15, and both are sold only to qualified professionals, so a person cannot buy either and take it at home. The clinical route costs what the examiner's hours cost: the University of South Florida Psychological Services Center, a training clinic, lists a Wechsler adult intelligence evaluation at 100 to 400 dollars on its clinic policies page, read on September 19, 2026, and private practices charge several times that.
Matrices tests are the narrow member of the family. Raven's Progressive Matrices, in the current Raven's 2 edition, present an incomplete figural pattern and ask which piece completes it, with almost no language in the instructions or the items, which is why the format is used where reading level or first language would contaminate a verbal test. The trade is breadth for portability: a matrices score estimates fluid reasoning well and says nothing about vocabulary, working memory or speed, as the page on Raven's 2 sets out.
ACIS belongs to this family by structure and to the last row of the table by delivery. It samples 20 subtests across six domains and reports a Full Scale IQ with a composite reliability of .9886 and a standard error of 1.60 points, against an adult reference frame of 3,243 English speaking records aged 16 to 90, but it is self administered in a browser rather than given by an examiner, which changes what the result can be used for. That difference is the subject of the disclosure that closes this page.
4 Achievement and Aptitude Tests: SAT, ACT and GRE
Admission tests are cognitive tests by the definition, and they differ from intelligence batteries in what they sample, whom they are scored against and what they are built to predict. The SAT, the ACT and the GRE present reading, writing and mathematics items drawn from what schools and colleges teach, under strict time limits, and report scores on scales set by the testing organization rather than on the IQ metric. Their reference group is the people who took the test in a given period, a self selected population of applicants, so a percentile on an admission test is a rank among candidates rather than a position in the general population of the same age. Their purpose is prediction: each is validated against grades in the first year of college or graduate school, not against a theory of cognitive ability.
The overlap with intelligence batteries is nevertheless large, because reasoning with learned material still draws on the general factor. Frey and Detterman, in Psychological Science in 2004, volume 15, issue 6, pages 373 to 378, reported a correlation of .82 between SAT scores and a general ability factor extracted from a broad test battery in a national longitudinal sample, and concluded that the SAT is mainly a test of general cognitive ability. The page on IQ versus the SAT works through what that correlation does and does not license, and the page on IQ versus the GRE applies the same analysis to the graduate test.
Three differences keep the families apart in practice. The first is content: an admission test rewards specific preparation because its items are drawn from a curriculum, while an intelligence battery is designed so that coaching on the items would invalidate the score, which is why its items are protected and its administrations restricted. The second is the reference group; a 90th percentile on the GRE is a rank among graduate applicants, a group already selected for ability, and cannot be converted to an IQ percentile by any table. The third is the report: admission tests deliver two or three section scores and a total, with no domain profile and no standard error printed for the individual, because the reader is an admissions office comparing thousands of candidates.
The word aptitude carries its own confusion. In the hiring literature an aptitude test is a narrow instrument for one ability, numerical, verbal, mechanical or spatial, chosen because it predicts performance in a specific role. In admissions the word once named the SAT, which was renamed away from it. In both senses an aptitude test is a cognitive test with a narrow sample and a specific criterion, and a person who wants to know their standing across domains is asking a question these tests were not built to answer.
5 Pre-Employment Cognitive Tests: Wonderlic and General Mental Ability
A pre-employment cognitive test is the shortest member of the family, and it exists because a few minutes of general mental ability measurement predict job performance better than most alternatives. The Wonderlic is the best known example, and its publisher's own candidate FAQ, read on September 19, 2026, describes the current product. The Wonderlic Select has three sections, one timed and two untimed, and takes about 35 minutes in total. The timed cognitive section contains 50 questions in 12 minutes, and the publisher states that test takers are not expected to finish. The untimed sections contain 58 motivational interest items and 150 personality items. Companies typically do not proctor it. The candidate receives a feedback report and no numerical cognitive score, and the publisher states that there is no universal good or average score because the three sections are combined against the demands of the specific job.
That last fact matters for anyone who arrives through a conversion chart. The page on Wonderlic scores and IQ explains that the charts refer to the classic Wonderlic Personnel Test, which reported a raw score from 0 to 50, and that a person tested in 2026 on the Wonderlic Select has no number to convert. The employer sees a fit report; the applicant sees a summary; nobody sees an IQ. The page on pre-employment cognitive tests covers the other products in the category, including the Criteria Cognitive Aptitude Test.
The evidence that justifies the family is meta-analytic. Schmidt and Hunter's review of 85 years of research in Psychological Bulletin in 1998, volume 124, issue 2, pages 262 to 274, put the validity of general mental ability tests for predicting job performance at .51, the highest of any single method they examined. Sackett, Zhang, Berry and Lievens revisited those estimates in the Journal of Applied Psychology in 2022, volume 107, issue 11, pages 2040 to 2068, and argued that the corrections for range restriction used across the earlier literature had been systematically too large. Their revised estimate for general mental ability is about .31, and structured interviews, at about .42, move to the top of the list. On either figure cognitive ability remains among the strongest predictors an employer can use; on the revised figure it is one strong predictor among several rather than the dominant one.
The market around these tests has its own vocabulary, and the phrase cognitive assessment platform is used for products that have little in common. The page on cognitive assessment platforms separates the hiring screens from the clinical, research and consumer products that share the label, and the page on the cost of cognitive testing platforms prices them. A candidate facing one of these tests is being measured against other applicants for that role, on a composite the employer has chosen, and the result is a decision rather than a profile.
6 Neuropsychological Batteries: Assessing Domains After Injury or Illness
A neuropsychological battery is a set of cognitive tests assembled around a clinical question, and the tests are one component of a process that ends in an interpretation. The American Academy of Clinical Neuropsychology's practice guidelines, published in The Clinical Neuropsychologist in 2007, volume 21, issue 2, pages 209 to 231, list the components: a referral question, a clinical interview, a review of records, standardized testing across cognitive domains, integration of the test data with the history and with behavioural observations, a written report, and feedback to the person tested. The testing alone usually takes several hours, and the page on neuropsychological testing for adults prices the whole process from provider fee pages and explains when it is the right tool.
The tradition the batteries follow is the domain by domain approach set out in Muriel Lezak's Neuropsychological Assessment, the standard textbook of the field: sample each function with instruments chosen for it and read the profile against the person's history. A typical adult battery covers general intellectual ability, usually with the WAIS-5 as its core; attention and processing speed; learning and memory in verbal and visual forms, with delayed recall; language; visuospatial and constructional skills, which is where block and card design tasks like the one in the image above appear; executive functions such as planning, inhibition and flexibility; motor and sensory function where relevant; and mood, because depression and anxiety produce cognitive complaints that the evaluation has to separate from a primary cause.
Two features distinguish this family from every other. The first is performance validity testing: the battery includes tasks whose purpose is to check that the other tasks were taken with full effort, and a report will state whether the scores are interpretable before it interprets them. No consumer test does this in the clinical sense. The second is the comparison being made. A neuropsychologist is often less interested in where a person stands relative to the population than in where they stand relative to their own expected level: a memory score that is average for the population but well below what the person's intelligence predicts is a finding, and co-normed instruments exist so that the comparison can be made with known error.
The examiner is a licensed psychologist with postdoctoral training in neuropsychology, and the instruments inside the battery are restricted to that qualification level, as the page on who can administer an IQ test explains. What the family answers is why: why a person's functioning has changed, whether a suspected disorder explains a pattern, whether a documented profile supports an accommodation. A person whose question is where they stand, with an error band, is asking for a different product and would overpay by an order of magnitude to get this one.
7 Brief Clinical Screens: the MMSE and the MoCA
A brief clinical screen is a cognitive test with a cutoff instead of a norm table, and it answers whether something may be wrong rather than how able a person is. The Mini-Mental State Examination was published by Folstein, Folstein and McHugh in the Journal of Psychiatric Research in 1975, volume 12, issue 3, pages 189 to 198, as a practical method for grading the cognitive state of patients at the bedside. It takes five to ten minutes, scores out of 30, and samples orientation to time and place, registration of three words, attention and calculation, recall of the three words, and language, including naming, repetition, following a command, reading, writing and copying a figure. It was designed to be given by a clinician without training in psychometrics, and it became the most widely used cognitive test in medicine for that reason.
The Montreal Cognitive Assessment was published by Nasreddine and colleagues in the Journal of the American Geriatrics Society in 2005, volume 53, issue 4, pages 695 to 699, as a screen for mild cognitive impairment, the stage the MMSE was known to miss. It is one page, takes about ten minutes, also scores out of 30, and adds tasks of executive function, visuospatial ability and higher level language that the MMSE lacks. In the original study of 94 people with mild cognitive impairment, 93 with mild Alzheimer disease and 90 healthy controls, a cutoff of 26 gave the MoCA a sensitivity of 90 percent for mild cognitive impairment and 100 percent for mild Alzheimer disease, with a specificity of 87 percent, where the MMSE detected 18 percent of the same mild cognitive impairment group. That is why 26 is the number attached to the test in public discussion, and why a total of 30 is a ceiling rather than an achievement.
What a screen cannot do follows from its design. It has a ceiling that most healthy adults reach, so it does not discriminate among people without impairment. It is scored against a cutoff rather than against an age and education distribution, though both tests are known to be affected by education and the MoCA adds a point for people with twelve or fewer years of schooling. It produces no IQ, no domain profile and no standard error. And a total below the cutoff is a flag for further evaluation, not a diagnosis, which the original papers say and which the neuropsychological battery above exists to settle. A screen used in public life as if it measured intelligence has been asked a question it was never built to answer.
8 Research Toolkits, School Group Tests and Online Batteries
Three further families share the definition, and each shows what a cognitive test looks like when it is built for a database, a classroom or a browser rather than for a clinic. The NIH Toolbox Cognition Battery is the research case. Weintraub and colleagues described it in Neurology in 2013, volume 80, issue 11, supplement 3, as a set of brief computerized measures of executive function, episodic memory, language, processing speed, working memory and attention, designed to be given in about 30 minutes to participants aged 3 to 85 so that studies across the lifespan could measure cognition on a common scale. Its tests include a flanker task for inhibitory control, a card sort for flexibility, list sorting for working memory, picture vocabulary for language and pattern comparison for speed. Scores are norm referenced and reported to the study rather than to the participant, and the toolkit is obtained by institutions rather than sold to individuals.
School group tests are the classroom case. The Cognitive Abilities Test, published by Riverside Insights, is described on its product page, read on September 19, 2026, as an assessment of verbal, quantitative and nonverbal reasoning for kindergarten through twelfth grade, used above all for identifying students for gifted and talented programs. It is given to whole classes by the school, with three batteries that map onto the verbal, quantitative and figural content of an intelligence battery, and it reports scores by age and by grade to the district rather than to the family. What it answers is a placement question for a school, on a scale of its own, and a CogAT result is not an IQ even though the two correlate.
Online self administered batteries are the browser case, and they range from public domain research items to normed commercial products to quizzes with a number at the end. The difference among them is checkable before paying: whether the site publishes a norm sample with its composition, the reliability of the score it prints, the standard error that follows, its time limits and its retake rules. The page on free versus validated IQ tests applies that checklist to the market. An online battery that publishes those things is a cognitive test by the definition on this page; one that publishes none of them is a sample of behaviour scored against nothing, and its number is not a position on any scale.
What the three families have in common is that the person tested is rarely the reader of the result. The study reads the toolkit, the district reads the group test, and, among online products, the person reads their own report only when the product was built to be read that way, which is why an honest online report prints the percentile and the standard error of each index.
9 What a Good Test Publishes: Norms, Reliability, Error and Validity
A cognitive test is only as good as its documentation, and the four things a publisher owes the reader are the same for a clinical battery, a hiring screen and an online product. The first is the norm sample. The Standards require that the population the norms represent be clearly described, with the size of the sample, its composition by age, sex, education and region, the year of collection and the procedure by which raw scores were converted to standard scores. The page on how IQ scores are normed explains why composition matters more than size and why norms date. A sample of 3,000 adults drawn to match a population is a reference frame; a sample of 300,000 website visitors is a convenience sample, and the larger number is the weaker norm.
The second is reliability, the consistency of the score, reported as internal consistency across items, as agreement between two administrations, or as a model based coefficient for a composite. The third follows from the second by arithmetic: the standard error of measurement, the expected size of the difference between a score a person obtained and the score they would obtain on average across many administrations. Harvill's instructional module in Educational Measurement: Issues and Practice in 1991, volume 10, issue 2, pages 33 to 41, gives the relation: the error equals the standard deviation of the scale multiplied by the square root of one minus the reliability. On an IQ scale a reliability of .90 produces an error of 4.7 points and a reliability of .9886 produces an error of 1.6 points, which is our arithmetic on the standard formula, and a 95 percent confidence interval extends about twice the error either side of the score. The page on reliability and validity explains how to read both figures.
The fourth is validity evidence, which the Standards organize by source: test content, response processes, internal structure, relations to other variables and consequences of testing. For an intelligence battery the internal structure evidence is usually a confirmatory factor analysis showing that the subtests group into the domains the publisher claims and that a general factor runs through them; for a hiring test it is a correlation with job performance; for a screen it is sensitivity and specificity against a diagnosis. ACIS reports its factor model, its fit statistics and the reliability and error of every index in its technical manual.
A fifth obligation arises whenever a test crosses a language or a culture. The International Test Commission's Guidelines for Translating and Adapting Tests, in their second edition in the International Journal of Testing in 2018, volume 18, issue 2, pages 101 to 134, set out 18 guidelines across six stages, from the decision that adaptation is appropriate through development, empirical confirmation, administration, score scales and documentation. Translation is not adaptation: an item that is easy in one language can be hard in another, and a norm collected in one country does not transfer without evidence. A test that publishes all five of these things can be evaluated; one that publishes none is asking to be trusted.
10 How Scores Are Expressed: Scaled, Standard, Percentile and T
Every cognitive test converts a raw count into a position on some scale, and the scales differ enough that a number without its scale is not a result. The conversion runs through the norm table: a raw score is located in the distribution of the reference group and expressed as a distance from the mean in units of the group's standard deviation. Publishers then choose a convention for the centre and the spread, and the table lists the conventions a reader is likely to meet. The last column describes one position, a standard deviation above the mean, which is about the 84th percentile of a normal distribution.
Score type
Centre and spread
Where it is used
The same position, one standard deviation above the mean
Raw score
A count of items correct or points earned
Every test, before norming
Not comparable across tests or ages
Scaled score
Mean 10, standard deviation 3
Individual subtests in Wechsler scales and ACIS
13
Standard score on the IQ metric
Mean 100, standard deviation 15
Wechsler indices, Stanford-Binet 5, ACIS indices and Full Scale IQ
115
T score
Mean 50, standard deviation 10
Many neuropsychological measures and personality inventories
60
z score
Mean 0, standard deviation 1
Research reporting
1.0
Percentile rank
1 to 99
Reports of every kind
About the 84th
Stanine
1 to 9, mean 5, standard deviation 2
School group tests
7
Cutoff total
Points out of a fixed maximum
MMSE and MoCA, out of 30
Not applicable; 26 and above is the MoCA paper's normal range
Test specific scale
Set by the testing organization
SAT, ACT, GRE
A percentile among test takers, not among the population
Three consequences follow. The first is that the same person holds several numbers at once on a single battery: a scaled score of 13 on a subtest, an index of 115, a percentile of 84 and, in a research report, a z of 1.0, and none of them contradicts the others. The second is that a standard score is only interpretable with its standard deviation. Older Stanford-Binet editions used 16, so 116 on those scales marks the position that 115 marks on a Wechsler scale, and the page on IQ scores versus percentiles gives the conversion; the safe comparison across scales is the percentile. The third is that a cutoff total is not a standard score at all: a MoCA of 27 does not place a person in a distribution, it places them above a line drawn for a clinical purpose.
The percentile deserves one caution of its own: it is nonlinear in the tails, so that the five points between 100 and 105 move a person from the 50th to the 63rd percentile while the five points between 130 and 135 move them from the 98th to the 99th, which is why a report should print both numbers.
11 What a Cognitive Test Is Not
A cognitive test is a narrow instrument, and most misuse comes from asking it questions outside the definition. It is not a personality test. Personality inventories sample typical behaviour, how a person usually acts, by self report, while a cognitive test samples maximal performance, the best a person can do on a task under fixed conditions. The Wonderlic Select bundles a personality section beside its cognitive section, and the bundling does not make the personality items cognitive; they measure a different construct with a different method and a different kind of validity evidence.
It is not a diagnosis. A MoCA total below 26 is a flag; an index score two standard deviations below the mean is a finding; a diagnosis is a clinical judgment that integrates test data with history, records, observation and validity evidence, which is the process the AACN guidelines describe and which no score performs on its own. The same holds for the disorders people most often bring to a cognitive test: dementia, learning disorders and attention deficit disorder are defined by patterns of symptoms across settings and over time, and a test session samples one hour of performance under conditions designed to hold attention.
It is not a measure of character, creativity, wisdom, motivation or emotional skill. Those are separate constructs with separate measures, and the correlation of each with measured cognitive ability is modest, as the page on what IQ measures documents for the intelligence battery case. A person with a Full Scale IQ at the 95th percentile has reasoning and knowledge at the 95th percentile of the reference frame; whether they are diligent, original or kind is not in the number.
It is not a measure of the brain. Imaging measures structure and activity; a cognitive test measures behaviour, and the relation between the two is a research question rather than a reading. It is not a fixed property of the person either: a score is a measurement with error, it moves with conditions on the day, and raw performance changes with schooling and with age, so that the honest form of any score is an interval around a value at a date. And the families are not interchangeable. A Wonderlic fit report is not an IQ, a MoCA total is not an IQ, a GRE percentile is not a population percentile, and a CogAT composite is a school placement score. Each is a cognitive test by the definition, each answers the question it was built for, and the error is not in any of them but in the reader who converts one into another.
12 How a Session Runs, and What Moves a Score on the Day
A cognitive test is a controlled occasion, and the controls exist because performance on the day is not the same thing as ability. A clinical administration follows a script. The examiner reads standardized instructions, presents one or two practice items on each subtest so that a failure reflects the task rather than a misunderstanding, times the speeded tasks with a stopwatch, and applies start and discontinue rules so that a person is not asked items far below or far above their level. The WAIS-5 takes 45 minutes for the seven subtest Full Scale IQ and 60 minutes for the ten primary index subtests, on Pearson's figures, and a neuropsychological battery takes several hours. The page on the adult testing process describes a session from the client's side, and the page on how long an IQ test takes compares the families by duration.
Online administration keeps the same elements without the examiner: instructions on screen, practice items before each subtest, a visible timer, and completion rules enforced by the software. What it cannot supply is a second person in the room, which is why a normed online battery depends on integrity controls, retake limits and completion requirements to protect the meaning of the score. The page on how to prepare for an IQ test separates preparation that improves the measurement, such as being rested and undisturbed, from preparation that contaminates it, such as practising the items.
Four things move a score on the day. Sleep is the first: a night of short sleep depresses attention and speeded performance most and reasoning least, and the page on sleep and IQ reviews the evidence and its size. Anxiety is the second. Moran's meta-analysis in Psychological Bulletin in 2016, volume 142, issue 8, pages 831 to 864, pooled 177 samples with 22,061 people and found a mean effect of g equal to -0.334 between anxiety and working memory capacity, a small to moderate effect that is largest on tasks that load on working memory, as the page on anxiety and IQ explains. Interruption is the third, and it acts through the speeded tasks, which is why processing speed indices are the most sensitive to conditions and should be read beside the other indices rather than as a deficit.
Practice is the fourth, and it is the reason retesting is restricted everywhere. Estevis, Basso and Combs retested 54 healthy adults on the WAIS-IV after three or six months and reported in The Clinical Neuropsychologist in 2012, volume 26, issue 2, pages 239 to 254, that Full Scale IQ rose by about seven points on average, with the largest gains on perceptual and processing speed tasks and little difference between the two intervals. A second administration of the same instrument within a year is therefore not comparable to the first, and every clinical protocol and every serious online battery limits retakes for that reason.
13 Where ACIS Sits, and How to Choose a Test for Your Purpose
We sell a paid online cognitive assessment, so treat this section as a disclosure and check it against the rest of the page. ACIS is an online self administered battery of 20 subtests across six domains, verbal comprehension, fluid reasoning, quantitative reasoning, visual spatial ability, working memory and processing speed, aligned with the Cattell-Horn-Carroll broad abilities Gc, Gf, Gq, Gv, Gwm and Gs. The page on the CHC model explains the taxonomy, and the page on the six cognitive domains describes each index and the subtests beneath it: Vocabulary asks for typed definitions and has a reliability of .950, Matrix Reasoning .930, Digit Span .890, and Symbol Search and Coding are timed speed tasks. The Full Scale composite has a reliability of .9886, a general factor loading of .958 and a standard error of 1.60 points, its confirmatory factor model fits with a CFI of .9761, and the adult reference frame holds 3,243 English speaking records aged 16 to 90, with a technical analysis set of 2,750 complete records, all per the technical manual, version 1.4.
The forms are priced once. The Full Scale form, 20 subtests and six domains, costs 50 dollars and takes about 175 minutes in sessions across 30 days; the Optimized form, 13 subtests and five domains, costs 30 dollars and takes about 110 minutes; the Quick form, six subtests and three domains, costs 15 dollars and takes about 45 minutes. Five subtests can be taken free without a card, and a quality guarantee refunds the purchase within five days if the report does not deliver what checkout described. Every report prints the percentile and the standard error of every index, with public classification bands that run from 90 to 109 Average to 160 and above Profoundly Gifted.
The limits are the ones this page has applied to everyone else. ACIS is not accepted for Mensa admission, for school placement or for clinical determinations of any kind, and nothing on this site claims otherwise. It has no examiner in the room, no clinical interview and no performance validity testing in the clinical sense, so it cannot answer the why questions that a neuropsychological battery exists for. Its norms are English speaking adults, so it is not interpretable for children or for people tested in a language they do not use daily. The page on official IQ tests sets the clinical, society and online routes side by side for a reader who needs a document rather than a measurement.
Choosing follows from the reader of the result. If a school, a court, a disability board or a physician will read it, only the clinical families produce what they accept, and a university training clinic is the least expensive door. If an employer or an admissions office will read it, the test is theirs to choose. If a physician suspects impairment, the screen comes first and the battery follows. If the reader is you, the question is which measurement is honest, and the answer is the one that publishes its norms, its reliability, its error and its limits before you pay.
Every figure above is traceable to one of the following, and each is linked at the point where it is used. ACIS reliability, loading, error and price figures are from the technical manual, version 1.4, and the ACIS home page. Provider and publisher pages were read on September 19, 2026 and will change. The standard errors for reliabilities of .90 and .9886 are our arithmetic on the formula in Harvill's module and are labelled as such.
American Educational Research Association, American Psychological Association and National Council on Measurement in Education. Standards for Educational and Psychological Testing, 2014 edition. apa.org.
International Test Commission. ITC Guidelines for Translating and Adapting Tests (Second Edition). International Journal of Testing, 2018, volume 18, issue 2, pages 101 to 134.
Pearson. Wechsler Adult Intelligence Scale, Fifth Edition (WAIS-5), product page with publication date, age range, completion time, qualification level and kit prices. pearsonassessments.com, read September 19, 2026.
University of South Florida Psychological Services Center. Clinic policies and fees. usf.edu, read September 19, 2026.
Frey M C and Detterman D K. Scholastic assessment or g? The relationship between the Scholastic Assessment Test and general cognitive ability. Psychological Science, 2004, volume 15, issue 6, pages 373 to 378.
Wonderlic. Wonderlic Select test: frequently asked questions from candidates. wonderlic.com, read September 19, 2026.
Schmidt F L and Hunter J E. The validity and utility of selection methods in personnel psychology: Practical and theoretical implications of 85 years of research findings. Psychological Bulletin, 1998, volume 124, issue 2, pages 262 to 274.
Sackett P R, Zhang C, Berry C M and Lievens F. Revisiting meta-analytic estimates of validity in personnel selection: Addressing systematic overcorrection for restriction of range. Journal of Applied Psychology, 2022, volume 107, issue 11, pages 2040 to 2068.
Board of Directors, American Academy of Clinical Neuropsychology. American Academy of Clinical Neuropsychology (AACN) practice guidelines for neuropsychological assessment and consultation. The Clinical Neuropsychologist, 2007, volume 21, issue 2, pages 209 to 231.
Folstein M F, Folstein S E and McHugh P R. "Mini-mental state": A practical method for grading the cognitive state of patients for the clinician. Journal of Psychiatric Research, 1975, volume 12, issue 3, pages 189 to 198.
Nasreddine Z S and colleagues. The Montreal Cognitive Assessment, MoCA: A brief screening tool for mild cognitive impairment. Journal of the American Geriatrics Society, 2005, volume 53, issue 4, pages 695 to 699.
Weintraub S and colleagues. Cognition assessment using the NIH Toolbox. Neurology, 2013, volume 80, issue 11, supplement 3.
Moran T P. Anxiety and working memory capacity: A meta-analysis and narrative review. Psychological Bulletin, 2016, volume 142, issue 8, pages 831 to 864.
Estevis E, Basso M R and Combs D. Effects of practice on the Wechsler Adult Intelligence Scale-IV across 3- and 6-month intervals. The Clinical Neuropsychologist, 2012, volume 26, issue 2, pages 239 to 254.
15 Frequently Asked Questions
What is a cognitive test?
A cognitive test is a standardized procedure that samples mental performance, on tasks of reasoning, memory, attention, speed, language or knowledge, under fixed conditions of timing and instruction, and scores the sample against a reference group. The score is a position within that group, and it carries a standard error because the test is a sample.
Is a cognitive test the same as an IQ test?
Not always. An IQ test is a cognitive test whose composite is expressed on a scale with a mean of 100 and a standard deviation of 15 against an age based population sample. Dementia screens, hiring tests, admission tests and research toolkits are also cognitive tests, and none of them produces an IQ.
What does a cognitive test measure?
It measures performance on a defined set of tasks: verbal knowledge, fluid reasoning, quantitative reasoning, visual spatial ability, working memory and processing speed in an intelligence battery; orientation, recall and language in a clinical screen; general mental ability in a hiring test. Which domains are sampled depends on the family and the question it was built for.
What are the main types of cognitive tests?
Eight families cover the field: intelligence batteries such as the WAIS-5 and matrices tests such as Raven's 2, admission tests such as the SAT and GRE, pre-employment tests such as the Wonderlic, neuropsychological batteries, brief clinical screens such as the MMSE and MoCA, research toolkits such as the NIH Toolbox, school group tests such as the CogAT, and online batteries.
What is a cognitive test used for?
Each family serves one reader: a clinician deciding on a diagnosis or accommodation, an admissions office predicting grades, an employer screening applicants, a physician checking for impairment, a research study measuring cognition on a common scale, a school identifying students for programs, or a person who wants to know their own profile with a stated error.
What is a cognitive assessment, and is it different from a cognitive test?
In clinical usage an assessment is the whole process, including the referral question, interview, record review, testing, integration and report, while a test is one instrument inside it. In consumer usage the two words are interchangeable, and the phrase cognitive assessment platform is applied to hiring, clinical, research and consumer products that have little in common.
What is an example of a cognitive test?
The Wechsler Adult Intelligence Scale, Fifth Edition, published in 2024 for ages 16 to 90, is the standard clinical example. The Montreal Cognitive Assessment, a ten minute screen scored out of 30, is the standard medical example. The Wonderlic Select, with 50 questions in 12 minutes, is the standard hiring example. ACIS is an online example with 20 subtests.
What is the difference between the MMSE and the MoCA?
Both are brief screens scored out of 30 and taking about ten minutes. The MMSE of 1975 samples orientation, registration, attention, recall and language; the MoCA of 2005 adds executive, visuospatial and higher language tasks. In the original MoCA study a cutoff of 26 detected 90 percent of people with mild cognitive impairment, where the MMSE detected 18 percent.
What is a normal score on the MoCA?
The original 2005 paper proposed 26 or above out of 30 as the normal range, with one point added for people with twelve or fewer years of education. A total below 26 is a flag for further evaluation rather than a diagnosis, and a total of 30 is the ceiling of the instrument rather than a measure of high ability.
What does the WAIS-5 measure, and how long does it take?
The WAIS-5 samples verbal comprehension, visual spatial ability, fluid reasoning, working memory and processing speed and reports a Full Scale IQ with index scores. Pearson lists 45 minutes for the seven subtest Full Scale IQ and 60 minutes for the ten primary index subtests, an age range of 16 to 90, and a qualification level restricting it to trained professionals.
What is the Wonderlic, and does it give you an IQ?
The Wonderlic Select is a hiring assessment with a 12 minute cognitive section of 50 questions plus untimed motivation and personality sections, about 35 minutes in total. The publisher states that candidates receive no numerical cognitive score and that there is no universal good or average score, so a 2026 result cannot be converted to an IQ.
How well do cognitive tests predict job performance?
Schmidt and Hunter's 1998 review put the validity of general mental ability at .51, the highest of any single method. Sackett and colleagues revised the figure to about .31 in 2022 after correcting for systematic overcorrection of range restriction, placing structured interviews, at about .42, above it. On either figure cognitive ability remains among the strongest predictors available.
Is the SAT a cognitive test?
Yes, by the definition: it samples reasoning with learned material under fixed conditions and scores it against a reference group. Frey and Detterman reported a correlation of .82 between SAT scores and a general ability factor in 2004. It differs from an IQ test in its reference group, test takers rather than the population, and in its scale.
What is the NIH Toolbox Cognition Battery?
A set of brief computerized tests of executive function, episodic memory, language, processing speed, working memory and attention, described by Weintraub and colleagues in Neurology in 2013 and designed to be given in about 30 minutes to people aged 3 to 85 so that studies can measure cognition on a common normed scale. Institutions obtain it; individuals cannot buy it.
What should a good cognitive test publish?
Four things at minimum: a description of the norm sample with its size, composition and year; the reliability of each score it reports; the standard error of measurement that follows from that reliability; and validity evidence fitted to the intended use, such as a factor analysis, a correlation with a criterion or sensitivity against a diagnosis.
What is the difference between a scaled score, a standard score and a T score?
Three conventions for one position. A scaled score has a mean of 10 and a standard deviation of 3, a standard score on the IQ metric a mean of 100 and a standard deviation of 15, and a T score a mean of 50 and a standard deviation of 10. So 13, 115 and 60 mark the same position.
Can a cognitive test diagnose dementia or ADHD by itself?
No. A screen total below a cutoff or an index well below the mean is a finding that calls for evaluation, and a diagnosis is a clinical judgment that integrates test data with history, records, observation and validity evidence. Dementia and attention deficit disorder are defined by patterns of symptoms over time, which one test session cannot observe.
What affects a cognitive test score on the day?
Sleep loss, which depresses attention and speeded performance most; anxiety, which a 2016 meta-analysis of 22,061 people linked to lower working memory capacity with an effect near a third of a standard deviation; interruptions, which act through the speeded tasks; and practice, which raised WAIS-IV Full Scale scores by about seven points at retest within six months.
Can you prepare for a cognitive test?
You can prepare the conditions: sleep, a quiet room, familiarity with the formats through the practice items every proper test provides, and a clear understanding of time limits. You cannot legitimately prepare the content: practising the items of an intelligence battery invalidates the score, and admission tests are the only family in which content preparation is expected.
Who can administer a cognitive test?
It depends on the family. Intelligence batteries and neuropsychological instruments are restricted to licensed psychologists with doctoral training, which publishers enforce through qualification levels. Screens such as the MMSE and MoCA are given by physicians, nurses and psychologists. Hiring and admission tests are administered by the organization. Online batteries are self administered under the rules of the software.
Is an online cognitive test like ACIS a real cognitive test?
It is a cognitive test by the definition on this page if it publishes a norm sample, reliability and standard error and administers items under fixed conditions. ACIS publishes those figures for 20 subtests against 3,243 adult records. What no online battery is, is a clinical instrument: ACIS is not accepted for Mensa admission, school placement or clinical determinations.
Take the assessment
You get a profile, not a number
ACIS measures six CHC domains across 20 subtests and reports each one with its own normed score and confidence interval, so you can see where you are strong and where you are not.