Intelligence measures for online studies, from 16 items to a full normed battery
A researcher adding a cognitive measure to an online study chooses among a 16 item public domain test, a 20 item matrix test, a nine item Raven form, commercial brief scales, tablet batteries and task libraries. This page compares them by what their own sources document: length, reliability, licence, cost and norms. ACIS sells one of the options, and it is described with the same limits.
Paper answer sheets like this one are being replaced by online delivery, yet the choice of measure still decides what the score can support.
0 The short answer
For an online study that needs intelligence only as a covariate, a short reasoning measure with published reliability and no licence fee, such as the 16 item ICAR or the 20 item Hagen Matrices Test, is usually the defensible choice, while a normed battery is justified only when the study needs population referenced scores or a profile across several abilities. The ICAR paper reports a Cronbach alpha of 0.81 for its 16 item composite and the Hagen Matrices Test documentation reports a Kuder-Richardson reliability of 0.78, but neither reports population norms in the documents we opened. The options that do (the NIH Toolbox, Creyos, ACIS) cost more in minutes, money or examiner qualification. The decisive distinction is between a score that adjusts one analysis and a score that describes a person against a reference population.
DisclosureACIS sells a competing paid cognitive assessment, and ACIS Professional is one of the ten options below. Every other instrument is described from a paper, manual or vendor page we opened, with the date it was read. ACIS facts come only from the home page as read on October 6, 2026, and no ACIS reliability, norm construction or psychometric statistic is quoted here. Where a source is silent, this page reports the absence and does not fill it.
0.81
Cronbach alpha of the 16 item ICAR composite in Condon and Revelle (2014), against 0.93 for the 60 item composite.
0.78
Kuder-Richardson reliability reported in the Hagen Matrices Test documentation, with a retest correlation of 0.75.
9 items
The Bilker and colleagues (2012) Raven forms, whose validation correlations with the 60 item total were 0.9063 and 0.8978.
1 Which Question Is the Intelligence Measure Answering?
The right measure depends on what the score will do in your analysis, and four jobs cover most of the reasons a team adds one to an online study. The first job is a covariate: the study is about attitudes, health behavior, decisions or language, and general cognitive ability is entered to adjust for a confound. The second is a screen or inclusion rule, where a score above or below a line decides who stays. The third is an outcome or profile, where cognition itself is the subject. The fourth is sample description, where the paper must say how the participants compare with a reference population. Each job asks a different thing of the score, and a measure that is excellent at one can be unusable at another.
Job
What the score must do
What matters most
Where a short free measure fails
Covariate
Rank participants within your own sample
Reliability, same form for everyone, few minutes
Low reliability leaves part of the confound uncontrolled
Screen or cut
Place a person above or below a line
Norms, precision near the line
A cut score on a raw count has no population meaning
Outcome or profile
Separate abilities such as Gf, Gc or Gwm
Several tasks, enough items per domain
One reasoning test cannot describe a profile
Sample description
State where the sample sits against a population
Norms from a stated reference group
A raw sum cannot be placed on the standard scale
Minutes are the budget that most designs ignore. Every minute a measure takes is paid for once per participant, so it scales with your sample. A 45 minute battery given to 1,000 people consumes 750 hours of participant time, and a 25 minute test consumes about 417 hours, our arithmetic from the durations. A short test is a different tool from a long one: it trades precision for the participant's time.
Every row of the comparison uses the same six checks: how many items and minutes the measure takes, what construct it targets, what reliability its source reports, whether norms exist, what licence or fee applies, and what evidence exists from unsupervised online use. A figure appears only if it was in a paper, manual or vendor page that we opened, and prices carry the date we read them. Where a vendor page lists no price, the table says so and does not estimate one. Instruments we could not document from a source we opened are left out rather than described from memory.
Reliability deserves one piece of arithmetic that changes how the numbers read. A reliability coefficient converts to a standard error of measurement through the standard deviation of the score and the square root of one minus the reliability. If a score were placed on the familiar scale with a standard deviation of 15, a reliability of 0.81 would imply a standard error of about 6.5 points, a reliability of 0.78 about 7 points, and a reliability of 0.93 about 4 points. These are our own illustrations, not figures from any source, and they ignore the fact that the instruments do not report on that scale. The page on reliability and validity explains the underlying logic, and the page on the standard deviation of 15 explains the scale.
The same coefficient also limits a covariate. A measure with reliability 0.81 can correlate with a perfectly measured criterion at no more than the square root of 0.81, which is 0.90, and a measure with reliability 0.78 at no more than about 0.88, again our arithmetic. In an adjustment, that gap is the part of general ability the covariate does not remove. For an exploratory study the loss is small. For a study whose conclusion is that an effect survives control for cognitive ability, it is the size of the doubt.
Validity evidence travels less well than reliability. A measure validated in supervised students has shown a relation in that group, not in crowdsourced participants on phones, so each section flags which sources studied unsupervised online use. The page on what IQ measures and the page on the CHC model define the constructs (Gf, Gc, Gq, Gv, Gwm, Gs) used here.
3 The ICAR16 and ICAR60: Public Domain Items Built for Online Use
The International Cognitive Ability Resource is the best documented of the free options we found for unsupervised online use, and its own paper is also the most candid about the price of being free.Condon and Revelle (2014) describe it in Intelligence as a public-domain measure built because the field lacked one for large-scale, remote data collection. Their abstract names the cost openly: the lack of copyright protection poses a theoretical threat to test validity, and the size of that threat is unknown. The paper then tests it on an online sample. In the version of the paper posted on the project site, that sample is 96,958 people from 199 countries, each given 12 to 16 items from a 60 item pool through a matrix sampling design, untimed and unsupervised.
The 60 items cover four formats: letter and number series (9 items), matrix reasoning (11 items), verbal reasoning (16 items) and three dimensional rotation (24 items). The 16 item ICAR Sample Test, the ICAR16, takes four items from each format, chosen as a representative set by difficulty and factor loading. Reliability is reported for each format and for the composites.
ICAR scale in Condon and Revelle (2014)
Items
Cronbach alpha
Omega total
Omega hierarchical
ICAR60 composite
60
0.93
0.94
0.61
ICAR16 composite
16
0.81
0.83
0.66
Letter and number series
9
0.77
0.80
0.66
Matrix reasoning
11
0.68
0.71
0.58
Verbal reasoning
16
0.76
0.77
0.64
Three dimensional rotation
24
0.93
0.94
0.78
Two features of the table matter for design. First, the 16 item composite is adequate rather than strong, and the authors themselves describe the alpha of 0.81 as adequate and the 60 item alpha of 0.93 as good. Second, the general factor saturation of the short form (omega hierarchical 0.66) is slightly higher than that of the long form (0.61), which means the short form is not a diluted version of the same signal, but a substantial share of its variance is still specific to its four formats. Matrix reasoning, the format closest to the classic Raven design, is the least reliable of the four, at 11 items.
Validity evidence in the paper is moderate and reported with its corrections. After correcting for the unreliability of self-reported scores, the ICAR16 correlated 0.59 with combined SAT scores and 0.52 with the ACT composite. In a university sample the ICAR16 was given alongside the Shipley-2, a brief commercial measure. The uncorrected correlations with the Shipley-2 composites were 0.41; corrected for restricted range they were 0.68; corrected for range and reliability they were 0.82 and 0.81. The authors note that these corrected values are comparable to the correlations the Shipley-2 itself shows with other ability measures. A researcher should read the uncorrected figure as what a student sample delivers in practice, and the corrected figures as an estimate of the relation in a broader population.
The measure has since been used outside its original site. In the Journal of Intelligence, Young and colleagues (2025) evaluated two ICAR formats (Puzzle Completion and Block Rotation) implemented in the Mobile Toolbox with 100 adults aged 18 to 82. They reported acceptable reliability, a correlation of 0.40 between Puzzle Completion and Raven's Progressive Matrices, 0.46 between Block Rotation and a mental rotation test, and non-significant practice effects. Those are formats from the ICAR family and not the ICAR16 composite, so they support the family rather than any one cut of it. A children's adaptation, the Ch-ICAR, was built for ages 11 to 14 because its authors wanted an alternative to proprietary Wechsler tests, which they say carry unfeasible administration times and financial costs.
On licence, read the project's own page. The ICAR project page at the Technical University of Dortmund, read on October 7, 2026, states that the ICAR is intended for academic use exclusively, lists a 60 item ICAR with a 12 item sample test, and points to an archive at the Leibniz Institute for Psychology. It does not name a licence type. "Public domain" in the paper and "academic use" on the project page are not identical statements, and a 12 item sample test is not the 16 item Sample Test of the 2014 paper, so confirm which version you are citing and which terms apply, especially for a commercial or non-academic study. The items are also public, which is the exposure risk the paper acknowledges. The page on the ICAR test and its limits covers interpretation for people who take it, including its norms. This page treats the ICAR as a research tool.
4 The Hagen Matrices Test: 20 Items, Free for Non-Commercial Research
The Hagen Matrices Test is a figural matrices test documented in an open access instrument database, and its documentation is unusually plain about where it should and should not be used. The entry in the GESIS collection of social science instruments, written by Heydasch and Schnaedter (2020) as an excerpt of Heydasch's 2014 dissertation material, describes a web-based test of induction and fluid reasoning in the CHC framework. It has 20 items, each a 3 by 3 incomplete matrix with a two minute limit per task, and exists in German, English and Spanish. The documentation gives the processing time as about 25 minutes, five for instructions and about 20 for the test, with a measured mean of 24.4 minutes and a standard deviation of 12.60.
The reliability figures are modest and stated without embellishment: a Kuder-Richardson reliability of 0.78 and a test-retest correlation of 0.75 over a mean interval of 78 days. The item analysis, on 1,339 participants, shows a mean difficulty of 0.37, with items ranging from 0.10 to 0.88, item-total correlations from 0.19 to 0.50, and a strong relation between position and difficulty. The first two items are very easy and seven are strongly difficult. The documentation reports convergent, divergent and criterion validity for the German version, with the highest correlations against the reasoning measures of the I-S-T 2000 R, a German intelligence test.
The documentation sets limits that many vendors would not print. It recommends the test for non-clinical adults aged 18 to 65, preferably with above average cognitive abilities, and warns that floor effects may arise, with frustration, in samples of lower cognitive ability. It was validated in samples aged 17 to 57. It should be used for group comparisons and correlative studies with sample sizes above 50, and it is not for individual diagnostics. The quality criteria were checked only for the German version, so an English or Spanish administration carries translation risk that the documentation does not resolve. The page on how IQ scores are normed explains why a score without a reference population cannot be read as an IQ.
On cost, the documentation states that the aim was a web-based, economical, valid alternative that is free of cost. The excerpt carries a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 licence, and the test itself can be requested by email from the author for non-commercial research. That is a request process, not a download, so build a few days into the schedule and keep the permission on file. Compared with the ICAR16, the Hagen test is longer, timed per item, documented as validated in one language, and recommended for samples with above average ability. In exchange it measures one construct, induction, and it documents a retest correlation of 0.75 over about eleven weeks, our conversion of 78 days.
5 Short Raven Forms: What the Nine Item Versions Do and Do Not Establish
The nine item Raven forms predict the full 60 item score closely, but predicting a total is a narrower claim than measuring intelligence, and the forms carry no norms of their own.Bilker and colleagues (2012), in Assessment, started from the 60 item Raven's Standard Progressive Matrices, which their abstract describes as a nonverbal estimate of fluid intelligence often used in clinical batteries and in research on patients with cognitive deficits. They used a Poisson predictive model to find a small subset of items that predicts the total score, and a second non-overlapping subset for repeated administrations.
Using nine items as predictors, the correlations with the full score were 0.9836 and 0.9782 in the fitting data for the two forms, and 0.9063 and 0.8978 in the validation data. The authors conclude that nine items predict the 60 item total with good accuracy, and that the two forms save 75 percent of administration time compared with a published 30 item form while achieving similar item and test level characteristics. The gap between the fitting and validation figures is the reason to quote the lower pair. A correlation of about 0.90 with the total means roughly 81 percent shared variance, our arithmetic from the correlation.
Three limits apply. First, the target of the paper is the 60 item total, so the forms inherit whatever the full test measures and whatever norms the full test has; they do not have their own. Second, the context in the abstract is clinical research, and the paper does not show how the forms behave in unsupervised crowdsourced samples. Third, the items belong to a published test, and the abstract says nothing about licence, so permission from the test's publisher is a question to settle before use. We did not verify the publisher's terms and this page does not state them. The page on Raven's 2 describes the current Raven product, and the page on fluid intelligence tests explains what matrix tasks measure and where they stop.
For a researcher who wants a matrices task online, the practical comparison is with the other short matrix instruments. The ICAR includes matrix reasoning items whose stimuli the paper describes as similar to those of the Raven, with an alpha of 0.68 on 11 items. The Hagen test is entirely matrices, with 20 items and an alpha of 0.78. A nine item Raven form is shorter than either and has the strongest published link to a long, well known test, but in the abstract we read it has no evidence from unsupervised online samples.
6 Vocabulary Measures: Wordsum and the Shipley-2
Vocabulary tests target crystallized ability (Gc), which makes them a different covariate from matrices, and their documentation shows how much item selection matters in a ten item test. The Wordsum is the ten item vocabulary test of the General Social Survey. Cor, Haertel, Krosnick and Malhotra (2012) analyzed it with item response theory and found that its items are very difficult or very easy, with the moderately difficult ones missing. Adding four moderately difficult items to make a 14 item battery, which they called Wordsumplus, outperformed the original on classical test theory indicators, reduced the standard error of ability estimates in the middle of the range, and showed higher concurrent validity, using national samples of thousands of American adults. Their recommendation is to use all 14 items.
The lesson for a researcher is that a short test can be reliable for the wrong part of the range. A ten item vocabulary test made mostly of very easy and very hard items separates the top and bottom of a sample well and the middle poorly, and most online samples live in the middle. It also measures acquired knowledge of word meaning, so it is sensitive to language background and schooling in a way that a matrix task is designed not to be. The page on verbal IQ tests covers what verbal measures capture.
The Shipley-2 is the commercial counterpart. The publisher listing at PAR, read on October 7, 2026, gives an administration time of 20 to 25 minutes, ages 7 to 89, three scales (Vocabulary for crystallized knowledge, Abstraction for fluid reasoning, and Block Patterns as a fluid alternative), Qualification Level B and a complete kit price of 293 dollars. It describes the test as a way to obtain quick ability estimates, screen for cognitive dysfunction or qualify participants for research studies, with normative data on 2,826 individuals and standard scores and percentiles. In the ICAR paper, university students completed it in 15 to 25 minutes, and Condon and Revelle cite its manual for correlations of 0.86 and 0.85 between its composites and the Full Scale IQ of the Wechsler Adult Intelligence Scale.
What the listing does not settle is delivery. A physical kit with a qualification level is built for supervised use, and the page we opened gave no terms for presenting its items in an unsupervised web survey. Before planning around it, ask the publisher whether and how its items may be delivered online. Its listing pairs published norms with a fee and a qualification level, which is a useful reminder that norms usually arrive together with both.
7 The Wonderlic and the Cognitive Reflection Test: Built for Other Jobs
The Wonderlic was built to screen job applicants and the Cognitive Reflection Test was built to study decision making, so neither is a drop-in intelligence covariate, and each needs a specific reason to be chosen. The Wonderlic Select page, read on October 7, 2026, states that its cognitive ability test consists of 50 questions answered in 12 minutes, and it describes the test as a way for employers to see whether applicants can learn from experience, solve problems and comprehend complex ideas, designed to predict job performance. The page lists no price and directs visitors to request a demo. That is 720 seconds for 50 questions, an average of 14.4 seconds each, our arithmetic, so the format rewards speed as much as depth. Its research use is not addressed on the page we opened, and a research licence is a question for the vendor.
The Wonderlic does appear in the research literature as a comparison point. Condon and Revelle report, citing the Shipley-2 manual, correlations of 0.64 and 0.60 between the two Shipley-2 composites and the Wonderlic Personnel Test, lower than the 0.86 and 0.85 the same composites show with the WAIS Full Scale IQ. A researcher who wants a quick score for a study would not obtain norms, a documented research licence and a published reliability from the vendor page alone. The page on how a Wonderlic score converts to IQ explains the conversion problem, and the page on pre-employment cognitive tests covers the hiring use that this page does not.
The Cognitive Reflection Test is a different animal. Frederick (2005), in the Journal of Economic Perspectives, introduced it as a three item test of cognitive reflection and showed that scores predicted patterns of decision making under risk and time. Toplak, West and Stanovich (2011) tested it against measures of cognitive ability, thinking dispositions and executive functioning in Memory & Cognition. They report that the CRT has a substantial correlation with cognitive ability and yet was a unique predictor of performance on heuristics-and-biases tasks, accounting for additional variance after the other individual difference measures were controlled. Their conclusion is that it measures a tendency toward miserly processing that intelligence tests do not capture.
That is a reason to include the CRT when your hypothesis concerns reflective thinking, and a reason not to use it as a stand-in for general ability. A three item test also has little room for reliability: three items give a score from 0 to 3, which cannot separate many participants. Public exposure is the other problem for a test this short, because online participants may have met it before. The page on the Cognitive Reflection Test documents that exposure for crowdsourced samples and walks through the test, so this page does not repeat it. If the aim is simply to control for ability, a matrix or series measure is the better covariate, and the CRT can sit beside it as a separate construct.
8 The NIH Toolbox and Creyos: Normed Batteries and What They Cost in Setup
The NIH Toolbox and Creyos are two of the options in this comparison that come with norms, and both ask for more setup than a free test: an examiner and a tablet for the first, a vendor quotation for the second. The NIH Toolbox Cognition Battery was described in Weintraub and colleagues (2013) in Neurology. Seven measures were designed to tap executive function, episodic memory, language, processing speed, working memory and attention. They were validated in English in 476 participants aged 3 to 85, with results on test-retest reliability, age effects and convergent and discriminant validity. The authors describe it as a brief, convenient set of measures to supplement other outcome measures in epidemiologic and longitudinal research and clinical trials, and say its computerized format and national standardization are meant to give researchers a common currency across studies.
Access is the constraint. The NIH Toolbox site, read on October 7, 2026, lists annual subscription prices for the version 3 app of 599.99 dollars for one or two devices, 1,499.99 dollars for three to six, and 2,499.99 dollars for seven to ten, for a full year of unlimited administrations. It states that administration is on an iPad and that the cognition tests require C-level qualifications or supervision by someone who holds them. Those are in-person conditions, and they explain why the Toolbox is a tool for studies where participants come to a site, not for a link sent to a crowdsourced panel. The page on who can administer an IQ test explains the qualification levels. The remote sibling is the Mobile Toolbox, in which the ICAR formats described earlier were implemented for remote self-administration, as Young and colleagues report.
Creyos, whose earlier name was Cambridge Brain Sciences, is a web platform. Its research page, read on October 7, 2026, states that the full assessment has 12 cognitive tasks and takes 35 to 45 minutes, that participants complete it online, that it draws on a normative database of more than 85,000 participants, and that its tasks have been used in more than 400 peer reviewed studies. It claims results comparable to traditional two to three hour pen and paper tests. These are the vendor's statements, and we did not verify them against the studies. The page lists no price and does not say what formats the data export takes. The page on the Creyos review examines the product in more detail.
Creyos offers twelve tasks and a vendor norm base, which is more than any free measure offers and less than a transparent technical document would. For a study that needs a multi-domain profile and cannot run a supervised session, it is a legitimate option to request a quotation for. Ask for the technical report, the norm sample description, the reliability of each task, the export format and whether participant-level data can be exported. Those questions apply to every vendor on this page, including ours.
9 Gorilla and Pavlovia: Task Libraries Are Delivery, Not Measures
Gorilla and Pavlovia deliver whatever task you build or import, so choosing them answers where the task runs and says nothing about whether the task measures intelligence.Anwyl-Irvine and colleagues (2020) presented Gorilla in Behavior Research Methods and demonstrated it with a simplified flanker task, replicating the conflict network effect in primary school children and adults, supervised and unsupervised, on participants' own computers and on computers supplied by the researcher, and over home internet and mobile connections. That shows reaction-time experiments can run reliably on the platform. It does not show that any intelligence task built on it is valid.
Bridges and colleagues (2020) compared timing across experiment software in PeerJ. Online studies did not deliver the same precision as lab systems, with slightly more variability in all measurements, though PsychoPy and Gorilla, broadly the best performers, achieved close to millisecond precision on a number of browser and system configurations. Timing precision matters for speeded and reaction time tasks and much less for untimed matrix or series items. Device matters too. Passell and colleagues (2021) examined 59,587 and 3,818 visitors to the TestMyBrain platform and found that users of mobile devices, particularly Android smartphones, showed slower measured reaction time than laptop and desktop users, after controlling for age, gender, education and an untimed vocabulary test. Their abstract does not claim that untimed reasoning tasks are immune, and neither does this page. Processing speed (Gs) tasks are the ones to watch, and the page on processing speed explains why.
Prices differ in kind. The Gorilla pricing page, read on October 7, 2026, shows no plan prices. It invites visitors to book a call and notes that prices exclude value added tax and that dollar prices are in United States dollars. The PsychoPy pricing page, read on October 7, 2026, states that PsychoPy is free, that Pavlovia credits cost 0.26 British pounds per participant for individual researchers without a licence, and that annual university and charity licences run 2,000 pounds for licence-only access, 2,200 for the standard tier and 5,200 for the plus tier. At the per-participant credit, 200 participants cost 52 pounds and 1,000 cost 260 pounds, our arithmetic and per credit consumed.
The page on the cost of testing other people adds recruitment fees, and the platform comparison page covers contracts and data terms. A library lowers the cost of delivery to almost nothing, and in exchange the validation, the scoring, the norms and the item security all become your responsibility. If you build a matrices task from scratch, you have a new, unvalidated instrument.
10 Side by Side: Items, Minutes, Reliability, Licence and Cost
Set in one table, the options sort into three groups: free reasoning tests that are short and unnormed, commercial scales with norms and fees, and platforms that sell delivery or batteries. The first table lists what each source documents. A cell that says not stated means the source we opened was silent, and silence is not evidence of absence.
Measure
Items and time
Construct in the source
Reliability in the source
Norms and licence in the source
ICAR16
16 items, untimed in the paper
Series, matrices, verbal reasoning, rotation
Alpha 0.81
No norm table in the paper; public domain in the paper, academic use on the project page
ICAR60
60 items, untimed in the paper
The same four formats
Alpha 0.93
The same
Hagen Matrices Test
20 items, about 25 minutes
Induction, fluid reasoning (Gf)
Kuder-Richardson 0.78, retest 0.75
None stated; free for non-commercial research on request
Nine item Raven forms
9 items per form, time not stated in the abstract
Fluid reasoning through the 60 item total
Validation correlations with the total 0.9063 and 0.8978
None of their own; licence not stated in the abstract
Wordsum
10 items, 14 in Wordsumplus
Vocabulary (Gc)
Improvements reported for 14 items
National survey samples; licence not stated in the abstract
Wonderlic Select
50 questions, 12 minutes
Cognitive ability for hiring, per the vendor
Not on the page
Not on the page; no price on the page
Shipley-2
20 to 25 minutes
Vocabulary (Gc), abstraction (Gf), block patterns
Not quoted here
Normative data on 2,826 individuals; kit 293 dollars; Level B
Cognitive Reflection Test
3 items
Cognitive reflection, not ability
Not quoted here
None; licence not stated in the sources opened
NIH Toolbox Cognition
7 measures
Six cognitive subdomains
Test-retest reported, figures not quoted here
National standardization; 599.99 dollars a year for one or two iPads
Creyos Research
12 tasks, 35 to 45 minutes
Cognitive tasks, per the vendor
Not on the page
Vendor norm base of more than 85,000; no price on the page
Gorilla and Pavlovia
Not measures
Task delivery
Not applicable
Gorilla shows no prices; Pavlovia 0.26 British pounds per participant
ACIS Professional
Quick 6 subtests, about 45 minutes; Optimized 13 subtests, about 110 minutes; Full 20 subtests, about 175 minutes
Six CHC domains
See the technical manual; none quoted here
Adult norms for ages 16 to 90; one time prices of 15, 30 and 50 dollars read on October 6, 2026
Cost has two parts that the table cannot show together: the fee for the instrument and the participant time that every minute consumes. The next table puts the participant time at three sample sizes, our arithmetic from the stated durations.
Measure and duration
50 participants
200 participants
1,000 participants
Hagen Matrices Test, 25 minutes
20.8 hours
83.3 hours
416.7 hours
Creyos Research, 35 to 45 minutes
29.2 to 37.5 hours
116.7 to 150 hours
583.3 to 750 hours
ACIS Quick, about 45 minutes
37.5 hours
150 hours
750 hours
ACIS Optimized, about 110 minutes
91.7 hours
366.7 hours
1,833.3 hours
ACIS Full Scale, about 175 minutes
145.8 hours
583.3 hours
2,916.7 hours
The fees follow the published prices, again our arithmetic. The free tests cost nothing in licence fees at any sample size, though the Hagen test needs a permission request and the ICAR needs a terms check. Pavlovia credits at 0.26 British pounds come to 13 pounds for 50 participants, 52 for 200 and 260 for 1,000 per task run. The NIH Toolbox subscription is 599.99 dollars for the year for up to two iPads, but the examiner time for an in-person session is the real limit. Where a team bought ACIS at the consumer one time prices read on October 6, 2026, Quick at 15 dollars would total 750, 3,000 and 15,000 dollars for those three sample sizes, Optimized at 30 dollars would total 1,500, 6,000 and 30,000, and Full Scale at 50 dollars would total 2,500, 10,000 and 50,000. The ACIS research page states what a research team pays, and a buyer should read it before relying on the consumer arithmetic. Recruitment and participant payment sit on top of all of these, and the page on the cost of testing other people prices them with platform fees.
11 Which Measure for Which Design, and What a Short Measure Cannot Do
Choose by the job the score does, then by minutes, then by licence, and accept that a short measure answers a narrower question than its label suggests. For a covariate in an unsupervised study, start with a short reasoning test that has published reliability: the ICAR16 if the academic terms fit, the Hagen test if one language and an above average sample fit, and a nine item Raven form only if permission from the publisher is in hand. Add a vocabulary measure when education or language background is part of the confound, because matrices and vocabulary tap different broad abilities (Gf and Gc). If your participants come to a site and an examiner is available, the NIH Toolbox and a supervised brief scale such as the Shipley-2 become possible, at the cost of qualification and equipment.
For a multi-domain profile, a short free test is the wrong tool. The profile has to separate abilities such as fluid reasoning, working memory and processing speed, and that needs several tasks per domain. Creyos and ACIS are the two options here built for it online, and both take 35 minutes or more. For a screen with a cut score, prefer an instrument whose source describes a reference group, and read the interval around the score near the line. For sample description, report the raw score and the instrument by name and avoid IQ language unless a normed scale exists. The sibling page on reporting IQ in research covers the wording.
A short measure cannot do five things. It cannot classify an individual: the Hagen documentation says it is for group comparisons and not individual diagnostics. It cannot make a controlled-for-intelligence claim airtight, because a covariate with reliability near 0.80 leaves part of general ability uncontrolled. It cannot be compared across instruments: a raw sum on the ICAR16 and a raw sum on the Hagen test are on different scales. It cannot be trusted in a new language because evidence travels poorly. And it cannot guarantee unsupervised honesty. The ICAR paper states plainly that in unsupervised online testing there are no safeguards against external resources, including those available on the internet. Design answers to that, such as time limits, response screening and attention checks, belong to the sibling page on online testing data quality.
Adaptive designs can shorten a test by choosing items from earlier answers, and the sibling page on adaptive IQ tests explains the idea. A supervised short Wechsler form belongs to a study with a qualified examiner, and the sibling page on the WAIS-5 subtests describes its formats. Finally, preregister the measure: name the instrument, version, scoring rule and exclusions before the data arrive, so the choice cannot follow the result.
12 Where ACIS Sits, and What a Buyer Should Do
ACIS sells one of the options on this page, and for most online studies that need a covariate it is not the best fit, which is why the comparison above recommends shorter free measures first. ACIS is a self-administered online assessment of 20 subtests in six CHC domains, reported as a Full Scale IQ and six primary indices: verbal comprehension, fluid reasoning, quantitative reasoning, visual spatial, working memory and processing speed. Read on October 6, 2026 on the home page, it has three forms with one time prices and no subscription. Quick costs 15 dollars, covers 6 subtests in 3 domains and takes about 45 minutes. Optimized costs 30 dollars, covers 13 subtests in 5 domains and takes about 110 minutes. Full Scale costs 50 dollars, covers all 20 subtests in six domains and takes about 175 minutes. Breaks are allowed, a free trial needs no card, and the offer includes a 5 day quality guarantee and 30 days to complete. Prices can change.
The score report gives indices and a Full Scale IQ on the standard scale with a mean of 100 and a standard deviation of 15, scaled subtest scores from 1 to 19, percentiles and a 95 percent confidence interval. Adult norms cover ages 16 to 90. ACIS Professional is the channel for a research or organizational team to administer ACIS to its own participants, and the research page and the Professional page describe how that works and what it does not do. The evidence behind the scores is in the technical manual, and this page does not quote it. A reader who wants the evidence should read it there and judge it by the same checks used for every instrument above.
The limits are as plain as the strengths. ACIS is online and unsupervised. It is not a clinical or diagnostic instrument, and it is not for hiring, school accommodations or admission to high IQ societies. It is available in English only. At 45 minutes for its shortest form, it asks participants for more time than the Hagen test and for a fee that the ICAR and Hagen tests do not charge. It fits a study that needs a profile across several domains, scores on a standard scale with an interval, and a form a participant can complete at home within 30 days. It does not fit a design that needs a few minutes of reasoning as a covariate.
Before buying any measure, a research team should do six things:
Ask the vendor or author for the technical documentation and read what the reliability and norms are based on.
Confirm in writing that research use and online delivery are permitted.
Pilot the measure with a small sample, check the score distribution for floor and ceiling effects, and compute reliability in your own data.
Decide in advance how scores will be reported, using the guidance in the sibling page on reporting IQ in research.
Budget participant minutes as a cost, not an afterthought.
Preregister the instrument, version and exclusions.
These steps follow the logic of the Standards for Educational and Psychological Testing (AERA, APA and NCME, 2014), the joint publication of the American Educational Research Association, the American Psychological Association and the National Council on Measurement in Education, which are open access at the standards site. Their central idea, as we read it, is that validity belongs to a proposed interpretation and use of scores, not to a test name, so the evidence has to match the use. The American Psychological Association's Guidelines for Psychological Assessment and Evaluation, approved in March 2020, point the same way. Guideline 8 asks for test selection, scoring and administration to reflect the appropriate normative comparison, and notes that the age of the norms and the continued relevance of the construct matter. Guideline 15 states that the essential criteria for evaluating technology enhanced measures are reliability, validity and fairness, and encourages review of the validation evidence when legacy tests are adapted for electronic presentation. Those guidelines are written for psychologists, but a research team choosing an instrument faces the same questions.
Every figure on this page comes from one of the sources below, opened by us. Papers are identified by DOI, and vendor and project pages carry the date we read them. ACIS facts come from the ACIS home page as read on October 6, 2026.
Anwyl-Irvine, A. L., Massonnié, J., Flitton, A., Kirkham, N., and Evershed, J. K. (2020). Gorilla in our midst: An online behavioral experiment builder. Behavior Research Methods, 52(1), 388-407. doi.org/10.3758/s13428-019-01237-x
American Educational Research Association, American Psychological Association, and National Council on Measurement in Education (2014). Standards for Educational and Psychological Testing. standardsforedtesting.org
American Psychological Association (2020). APA Guidelines for Psychological Assessment and Evaluation, approved by the APA Council of Representatives, March 2020. apa.org
Bilker, W. B., Hansen, J. A., Brensinger, C. M., Richard, J., Gur, R. E., and Gur, R. C. (2012). Development of abbreviated nine-item forms of the Raven's standard progressive matrices test. Assessment, 19(3), 354-369. doi.org/10.1177/1073191112446655
Bridges, D., Pitiot, A., MacAskill, M. R., and Peirce, J. W. (2020). The timing mega-study: comparing a range of experiment generators, both lab-based and online. PeerJ, 8, e9414. doi.org/10.7717/peerj.9414
Condon, D. M., and Revelle, W. (2014). The international cognitive ability resource: Development and initial validation of a public-domain measure. Intelligence, 43, 52-64. doi.org/10.1016/j.intell.2014.01.004. Figures were read in the version posted on the ICAR project site.
Cor, M. K., Haertel, E., Krosnick, J. A., and Malhotra, N. (2012). Improving ability measurement in surveys by following the principles of IRT: The Wordsum vocabulary test in the General Social Survey. Social Science Research, 41(5), 1003-1016. doi.org/10.1016/j.ssresearch.2012.05.007
Dutry, M., Vereeck, A., Duyck, W., Derous, E., Schelfhout, S., Szmalec, A., Woumans, E., Schittekatte, M., Debeer, D., and Dirix, N. (2025). Validation of the Children's International Cognitive Ability Resource (Ch-ICAR). Behavior Research Methods, 57(2), article 66. doi.org/10.3758/s13428-024-02591-1
Frederick, S. (2005). Cognitive reflection and decision making. Journal of Economic Perspectives, 19(4), 25-42. doi.org/10.1257/089533005775196732
Heydasch, T., and Schnaedter, S. (2020). The Hagen Matrices Test (HMT). Zusammenstellung sozialwissenschaftlicher Items und Skalen (ZIS), GESIS. access.gesis.org/zis/655
Passell, E., Strong, R. W., Rutter, L. A., Kim, H., Scheuer, L., Martini, P., Grinspoon, L., and Germine, L. (2021). Cognitive test scores vary with choice of personal digital device. Behavior Research Methods, 53(6), 2544-2557. doi.org/10.3758/s13428-021-01597-3
Toplak, M. E., West, R. F., and Stanovich, K. E. (2011). The Cognitive Reflection Test as a predictor of performance on heuristics-and-biases tasks. Memory & Cognition, 39(7), 1275-1289. doi.org/10.3758/s13421-011-0104-1
Weintraub, S., Dikmen, S. S., Heaton, R. K., Tulsky, D. S., Zelazo, P. D., Bauer, P. J., and others (2013). Cognition assessment using the NIH Toolbox. Neurology, 80(11, Supplement 3), S54-S64. doi.org/10.1212/wnl.0b013e3182872ded
Young, S. R., Kim, J., McKee, K., Doyle, D. R., Novack, M. A., Revelle, W., Gershon, R., and Dworak, E. M. (2025). Validation of International Cognitive Ability Resource (ICAR) implemented in Mobile Toolbox (MTB). Journal of Intelligence, 13(12), 154. doi.org/10.3390/jintelligence13120154
What is the best intelligence measure for an online study?
There is no single best measure, because the right one matches the job the score does. For a covariate in a crowdsourced study, a short reasoning test with published reliability, such as the ICAR16 or the Hagen Matrices Test, usually fits. For population referenced scores or a multi-domain profile, a normed battery is needed.
What is the shortest valid IQ measure for a survey?
The shortest option with published evidence among the sources here is a nine item Raven form, whose scores correlated about 0.90 with the 60 item total in validation data. Shortness has costs: that evidence concerns predicting a total, not norms or unsupervised online use, and shorter forms leave more of the construct unmeasured.
Can I use the ICAR16 in my online study?
Yes for academic research, subject to the terms the project publishes. The ICAR paper presents it as a public-domain measure, while the project page states it is intended for academic use exclusively. Confirm the version, the terms and the archive source before collecting data, especially for any commercial or non-academic purpose.
Is the ICAR normed?
The 2014 paper describes the ICAR as a public-domain item pool and reports reliability and validity evidence, but the version we read does not present a table converting raw scores into IQ. Scores therefore describe position within your own sample unless you add an external reference.
Can I use Wordsum as an IQ measure?
Wordsum is a ten item vocabulary test from the General Social Survey, so it indexes crystallized ability (Gc) in word knowledge, not general intelligence as a whole. Its uneven item difficulty led researchers to propose a 14 item version. Treat it as a verbal covariate and do not label it an IQ.
How long should a cognitive task be in an online study?
Match the length to the job and the budget. A covariate needs minutes, not hours: the Hagen test documents about 25 minutes including instructions, and the ICAR16 has 16 untimed items. The multi-domain batteries here run from 35 to 175 minutes, and every added minute is paid for by every participant.
Do I need norms to include a cognitive measure as a covariate?
Not for adjusting within your own sample, where a reliable rank ordering is enough. You need norms when you must place participants against a population, apply a cut score, or describe the sample as average or above average. Without a reference group, a raw sum has no population meaning.
How reliable is the ICAR16?
In Condon and Revelle, the 16 item composite had a Cronbach alpha of 0.81 and an omega total of 0.83, which the authors call adequate. The 60 item composite reached 0.93. For a covariate that is workable, though it leaves a measurable share of general ability uncontrolled.
How reliable is the Hagen Matrices Test?
The documentation reports a Kuder-Richardson reliability of 0.78 and a test-retest correlation of 0.75 over a mean of 78 days for the 20 item German version. Its quality criteria were checked only for German, so translated versions carry added uncertainty that the documentation does not quantify.
Is a nine item Raven form as good as the full test?
It predicts the 60 item total well, with validation correlations of 0.9063 and 0.8978 for the two forms, but that is a statement about reproducing a total score. The forms carry no norms of their own, and the abstract reports no evidence from unsupervised online samples.
Can I use the Wonderlic in research?
The Wonderlic is marketed for hiring: its vendor page describes 50 questions in 12 minutes and lists no price or research terms. A research team should ask the vendor about research licensing, norms and reliability documentation, and should not treat a Wonderlic score as an IQ without a documented conversion.
Is the Cognitive Reflection Test an intelligence test?
No. It was designed to measure the tendency to override an intuitive answer and reflect. Research finds it correlates substantially with cognitive ability yet predicts heuristics-and-biases performance beyond it. Use it when your hypothesis concerns reflective thinking, and pair it with a separate ability measure for adjustment.
What does the NIH Toolbox cost?
The version 3 app subscription listed on its site on October 7, 2026 was 599.99 dollars a year for one or two devices, 1,499.99 for three to six and 2,499.99 for seven to ten. It runs on an iPad, and the cognition tests need C-level qualified administration or supervision.
Do Gorilla and Pavlovia include an intelligence test?
The sources here describe their timing and pricing, not an intelligence test with published reliability or norms. These platforms host tasks that researchers build or import. Choosing one decides where a task runs, while validity, scoring and item security remain the researcher's responsibility.
Does controlling for a short intelligence measure remove the confound?
Only partly. Adjusting for an imperfectly reliable measure leaves some confounding by general ability in place, because the covariate captures only its reliable share. With a reliability near 0.80, the covariate can correlate with true ability at no more than about 0.90, our arithmetic.
Can I use a free test to exclude participants by IQ?
A short free measure is a weak basis for exclusion. A cut score on a raw count has no population meaning without norms, and a score with reliability near 0.80 misclassifies people near the line. If exclusion by ability is required, pick a measure with a documented reference group and report its limits.
Should I report the score as an IQ?
Report the raw score and name the instrument, and reserve IQ language for scores on a normed scale. A sum of correct items is not an IQ. State the reliability you obtained in your sample and any exclusions in the same methods paragraph, so readers can judge the measure.
Can participants look up answers in an online study?
It is possible, and the ICAR paper says so: with unsupervised online testing there are no safeguards against external resources, including those on the internet. Item exposure, repeated attempts and outside help are managed with design choices and response screening, not prevented by the instrument itself.
Does ACIS fit a research study?
ACIS sells a paid assessment, and ACIS Professional is the route for administering it to a research team's participants. Its Quick form takes about 45 minutes, so in most covariate designs a shorter free measure fits better. It suits studies that need a profile across several domains, within the limits stated.
Can I translate a test into another language for my study?
Translation changes how items behave, so evidence from one language does not carry over automatically. The Hagen documentation says its quality criteria were checked only for the German version. If you translate, pilot the new version and report its reliability in your own sample.
What should I preregister about the cognitive measure?
Name the instrument and version, the number of items, time limit and delivery device, the scoring rule, how missing responses are handled, any exclusion based on the score, and the reliability you will report. Registering these before data collection prevents the measure from being chosen after the results.
Take the assessment
You get a profile, not a number
ACIS measures six CHC domains across 20 subtests and reports each one with its own normed score and confidence interval, so you can see where you are strong and where you are not.