Public Domain Assessment

The ICAR Test and the Percentile It Cannot Give You

The International Cognitive Ability Resource is a public domain cognitive ability battery built by academic researchers and used in hundreds of studies. Its reliability evidence is published and good. Its norms do not exist. That single gap explains why searching for an ICAR 60 score to IQ conversion returns nothing usable.

A flat illustration of the top half of a man's face with a glowing yellow lightbulb above his hair.
The ICAR publishes reliability and validity evidence for each of its item types, but it publishes no table converting a raw score into a standardized IQ metric.

0 The Quick Answer on the ICAR Test

The ICAR test is a free, public domain battery of cognitive ability items developed by academic psychometricians, and it measures general cognitive ability well enough for research, but it has never published a table that converts a raw ICAR score into an IQ or a percentile. Both statements are true at once, and the second one is the reason this page exists.

ICAR stands for the International Cognitive Ability Resource. David Condon and William Revelle introduced it in the journal Intelligence in 2014, in a paper titled The international cognitive ability resource: Development and initial validation of a public-domain measure. The argument of that paper was straightforward. Commercial cognitive tests are expensive, licensed, and legally restricted, which makes large scale online research awkward or impossible. A public domain item pool would remove all three obstacles. The paper then tested whether giving the items away destroyed their validity, and concluded that it did not.

What the paper did not do, and what nothing published since has done, is establish a reference population against which an individual raw score could be placed. The 96,958 people in the founding study were self selected internet volunteers. They were 66 percent female, their median age was 22, and 78.1 percent of them were in the United States. That is an excellent sample for estimating how items correlate with each other. It is not a sample from which anyone can honestly tell you that 41 correct out of 60 puts you at the 84th percentile of adults.

So the honest complication is this: the ICAR is better documented than most tests you can take online for money, and it still cannot tell you where you stand. Those are two different jobs, and the ICAR was only ever built to do one of them. It is free in the sense that matters to a research budget rather than the sense meant by the free online IQ tests aimed at consumers, which sell a score and give away the questions rather than the reverse.

60 items

The full ICAR 60 form: 9 letter and number series, 11 matrix reasoning, 16 verbal reasoning and 24 three dimensional rotation items (Condon and Revelle, 2014).

.93

Coefficient alpha for the 60 item form in the founding online sample of 96,958 participants (Condon and Revelle, 2014).

No norm table

No peer reviewed publication or ICAR project page converts an ICAR raw score into a standardized IQ metric or a population percentile.

1 What the International Cognitive Ability Resource Actually Is

The ICAR is not a test in the way that a commercial battery is a test. It is an item repository plus a research consortium, and the difference matters for everything that follows.

A commercial instrument is a fixed object. It has a manual, a publisher, a qualification level, a fixed administration order, a fixed number of items and a printed normative table with a date on it. The ICAR has none of those. It has item types, a growing pool of items within each type, and a set of papers describing how those items behave. Researchers assemble whatever form they need from the pool.

The project began with four item types validated in the 2014 paper. By the time Dworak, Revelle, Doebler and Condon wrote their overview for Personality and Individual Differences, the pool had grown to 19 types of content covering spatial, verbal, mathematical and perceptual ability, including two dimensional and three dimensional rotation, figural analogies, propositional reasoning, arithmetic, emotion recognition, melodic discrimination, a perceptual maze task and a situational judgment task. That same paper reports close to 1,000 ability items administered through the associated SAPA data collection project, roughly 1,500 registered users of the content repository, and citations in more than 100 academic articles between 2017 and 2019 alone. Funding came from a collaboration of the United States National Science Foundation, the German DFG and the United Kingdom ESRC.

Two details about access are worth stating precisely, because people get them wrong. First, the content is described by its authors as maintained in the public domain for non commercial research purposes. Second, the actual content repository requires user registration, and the ICAR project site directs prospective users to register rather than simply downloading the bank. The items are free. They are not indiscriminately posted.

What the ICAR is not is also worth listing. It is not a clinical instrument. It is not qualified to any professional level, because there is no publisher to qualify it. It produces no score report, so none of the components that a detailed results report is expected to carry, from index scores through percentiles to confidence intervals, exist here at all. It has no interpretive guidance for an individual test taker, by design, because the 2014 paper explicitly argued that interpretive feedback materials have relatively little value in research administration. If you want to understand what the broader category of cognitive ability testing looks like when a publisher does take responsibility for interpretation, that is a different class of product entirely. A publisher that takes that responsibility also controls who may buy the product: an individually administered battery sits at Level C, which asks for a doctorate plus formal training in ethical administration, scoring and interpretation before an order will be filled, and the guide to who clears each qualification tier shows how much of the field that leaves out.

The axis people usually reach for is free against paid, and it is the wrong axis. The useful contrast is free against validated and normed, and the ICAR occupies an unusual cell in that grid: it is free, it is validated, and it is unnormed. Almost nothing else sits there.

2 The Four Original Item Types

Everything most people call the ICAR is really four item types, and they are not equally good. Condon and Revelle validated Letter and Number Series, Matrix Reasoning, Verbal Reasoning, and Three Dimensional Rotation, and reported each one separately rather than hiding the weak ones inside a composite.

Letter and Number Series presents an alphanumeric series and asks what comes next. It is an induction task, closest in spirit to number series reasoning of the kind ACIS measures in its own number series subtest. Nine of these items are in the 60 item form.

Matrix Reasoning presents three by three arrays of geometric shapes with one of the nine cells missing, and asks which shape best completes the array. This is the format everyone recognizes, and it is the format used by the ACIS matrix reasoning subtest and by Raven's family of tests. Eleven items are in the 60 item form.

Verbal Reasoning mixes logic, vocabulary and general knowledge questions inside a single strand, which is a different construction from a dedicated word knowledge task such as the 45 question ACIS vocabulary subtest, where every item probes definitional knowledge and nothing else. Sixteen items are in the 60 item form.

Three Dimensional Rotation presents cube renderings, each face carrying a different image, and asks which response option is a possible rotation of a target cube. It is a spatial task in the sense that visual spatial subtests generally are. Twenty four items are in the 60 item form, making it by far the longest strand.

Item typeItems in ICAR 60AlphaOmega totalOmega hierarchical
Three Dimensional Rotation24.93.94.78
Verbal Reasoning16.76.77.64
Matrix Reasoning11.68.71.58
Letter and Number Series9.77.80.66
Full ICAR 6060.93.94.61

All values above come from Table 3 of Condon and Revelle (2014), computed on composites of Pearson correlations between items. Read the Matrix Reasoning row carefully. The authors themselves called its internal consistency marginally adequate and wrote that the 11 items were not uniformly measuring a singular latent construct. The psychTools R package documentation, maintained by Revelle, goes further: it notes that two of the matrix reasoning problems do not have monotonically increasing trace lines, so at moderately high ability the probability of a correct answer actually dips. That is an unusual thing for a test author to publish about his own items, and it is one of the reasons the ICAR deserves respect. Reporting each strand separately instead of burying the weak one inside a composite is precisely the practice that a breakdown of subtest types and their confounds argues for, and it is why a single ICAR total conceals more than it reveals. If you want the general framework these four types sit inside, see the cognitive domains model.

3 ICAR 16 and ICAR 60, the Two Standard Forms

The ICAR 16, also called the ICAR Sample Test, is four items from each of the four original types, drawn from the same 60 item pool. It exists so that a researcher who needs an ability covariate but not an ability study can spend a few minutes of participant time instead of half an hour.

The reliability trade is exactly what you would expect from item count, with one twist that is not what you would expect. Alpha falls from .93 on the long form to .81 on the short form, and omega total falls from .94 to .83. But omega hierarchical, the proportion of variance attributable to a single general factor, goes the other way: .61 for the ICAR 60 and .66 for the ICAR 16. The long form is more reliable and less purely general, because 24 of its 60 items are rotation items and that block carries a large group factor of its own. If your interest is a clean estimate of the general factor, the short form is not merely a cheaper substitute.

PropertyICAR 16 (Sample Test)ICAR 60
Items16, four per type60, unevenly distributed
Alpha.81.93
Omega total.83.94
Omega hierarchical.66.61
Suggested upper time limit16 minutes for adults aged 18 to 25Not published
Validated against an individually administered batteryYes, against WAIS-IV in one sample of 97No published study of this kind
Reference distribution for individual scoringNone publishedNone published

The 16 items are identified by name in public. They are reason.4, reason.16, reason.17 and reason.19; letter.7, letter.33, letter.34 and letter.58; matrix.45, matrix.46, matrix.47 and matrix.55; and rotate.3, rotate.4, rotate.6 and rotate.8. Anyone can install the psychTools package in R and load a dataset called ability containing scored responses from 1,525 SAPA participants on exactly those 16 items. That level of openness is the point of the project, and it is also, as a later section explains, the source of its one structural vulnerability.

Beyond the two named forms there is no canonical version. Dworak and colleagues report that published studies have used as few as four ICAR items and as many as 60, and that some studies drew on a single item type while others combined five. This flexibility is a genuine feature for research design. It is fatal for score comparability across studies, which is a separate matter from whether any individual study measured well. Breadth is also the thing that buys accuracy in the first place, which is the argument behind what makes an assessment comprehensive: sixteen items sampling four narrow tasks is a screener, and calling it a battery does not make it one.

4 Whether the ICAR Is Timed, and How Long It Takes

The ICAR is not timed. Timing is not part of administration and it is not part of scoring. This is deliberate, documented, and one of the clearer contrasts with the commercial batteries people compare it against.

Condon and Revelle listed three goals when they developed the first four item types. The second of those goals was to avoid timed items, because telemetric assessment over the open internet introduces technical variability that makes any timing measurement unreliable. A participant on a slow connection is not a slower thinker. Rather than measure something contaminated, they chose not to measure it at all, and the 2014 paper states plainly that none of the items were timed in those administrations.

The later overview by Dworak and colleagues is more practical about it. The ICAR measures are power tests, so timing plays no role in scoring, but the authors acknowledge that researchers still need administration to move along. Their guidance for the 16 item Sample Test is a maximum of 16 minutes for young adult participants aged 18 to 25, with the note that most people finish in about half that. They also warn that three dimensional rotation items take some participants much longer than the other types, so a form heavy in rotation will not behave like the Sample Test. In practice, published studies have ranged from imposing no limit at all to capping administration at 10 minutes. Those minutes sit at the short end of the whole category: Pearson states 45 minutes for the seven subtest WAIS-5 Full Scale and Stoelting 50 minutes for the ten subtest Stanford-Binet 5, figures collected in the comparison of stated administration times by instrument, so an ICAR 16 buys an ability covariate for about a third of the testing time a short clinical battery costs.

Two consequences follow, and both are frequently missed.

  • An ICAR score contains no information about processing speed. If your research question touches processing speed as a cognitive domain, the ICAR does not address it, and you cannot back into it from response latencies that were never designed to be scored. Measuring it requires a task built to be timed, of the kind the ACIS symbol search subtest uses, where rapid visual scanning under accuracy pressure is the scored quantity rather than an artifact.
  • An untimed, unproctored item is an item a participant can look up. Vocabulary was left out of the ICAR pool for exactly this reason, as the next sections describe. This is a general property of remote testing rather than an ICAR defect, and it is the same reason time limits on IQ tests exist in supervised settings.

None of this makes the ICAR worse than a timed test. Power tests and speeded tests measure overlapping but distinct things, and a research design that wants pure reasoning under no time pressure is better served by an untimed instrument. It does mean that comparing an ICAR result against a timed battery's score is comparing two different administrations, not two estimates of the same quantity.

5 What an ICAR Raw Score Actually Is

An ICAR score is a count of correct answers and nothing else. It is not scaled, not age adjusted, not standardized, and not anchored to any distribution. If you took the ICAR 60 and got 41, your score is 41. A raw count of that kind is not the sort of input that the conversion between a score and a percentile accepts, since that conversion starts from a standardized metric with an agreed mean and standard deviation. The interesting question is what 41 can be compared against, and the answer is less than most people assume.

Start with how the reference data were collected. The founding sample was gathered through Synthetic Aperture Personality Assessment, a matrix sampling procedure in which each participant receives a different overlapping subset of items. Condon and Revelle report that participants were administered subsets of 12 to 16 items, that the median number of administrations per item was 21,764, and that the median number of pairwise administrations for any two items was 2,610. This is a powerful design for estimating a covariance structure from a very large number of people. It also means that almost nobody in the founding dataset ever sat the full 60 item form, so there is no large empirical distribution of ICAR 60 total scores in the source data at all. The composite properties were estimated, not directly observed on complete forms.

Now consider precision on the short form. The 2014 paper reports the standard deviation of ICAR 16 scores in the unrestricted online sample as 1.86 raw points. One standard deviation on a 16 item test is therefore under two items. Combine that with the published alpha of .81 and the standard error of measurement works out to 1.86 multiplied by the square root of 1 minus .81, which is about 0.81 items. A conventional 95 percent confidence interval around an individual ICAR 16 score is therefore roughly plus or minus 1.6 items.

What that arithmetic meansTwo people whose ICAR 16 scores differ by one item are not distinguishable. Two people who differ by three items are barely distinguishable. This is not a flaw in the ICAR; it is what a 16 item test costs, and every short form pays it. The arithmetic above uses only figures published by Condon and Revelle, and it applies within their sample rather than to any general population.

For comparison of a different kind, Study 3 of the same paper administered the ICAR 16 to 137 students at a selective private university and found a standard deviation of 1.48 rather than 1.86. Same instrument, same 16 items, a fifth less spread, purely because of who was in the room. That is range restriction working exactly as textbooks say it does, and it is a preview of why a fixed conversion table would be misleading. If the standard deviation of the metric depends on the sample, then so does any percentile you might build on top of it. This is the same problem the 15 point standard deviation convention was invented to solve, and the ICAR never adopted a solution.

6 Why an ICAR Score Cannot Be Turned Into a Percentile

People searching for an ICAR 60 score to IQ conversion find nothing useful because nothing useful exists, and the reasons are structural rather than accidental. There are four of them, and each would be disqualifying on its own.

There is no reference population. A percentile is a statement about a defined group. The ICAR's large sample came from volunteers who chose to take a personality survey on the internet and were then offered ability items. Condon and Revelle describe the sample honestly: 96,958 people from 199 countries, 66 percent female, median self reported age 22, with 78.1 percent from the United States. Nobody drew that sample from a frame, stratified it, or weighted it. It tells you a great deal about item behavior and nothing defensible about population standing. The mechanics of what would be required instead are covered on the page on how IQ scores are normed.

There is no fixed form. A percentile table belongs to a specific test administered a specific way. The ICAR is a pool from which researchers build their own forms, and they do: four items in some studies, 60 in others, one item type in some, five in others. Two people who both say they took the ICAR may have taken almost nothing in common.

There is no stratification by age or education. Fluid reasoning peaks in early adulthood and then declines, while crystallized ability follows a different curve entirely, a divergence set out on the page contrasting fluid and crystallized ability. A norm that ignores it will misplace almost everyone over forty. Any usable adult norm therefore has to be conditioned on age in the way age banded average scores are, and no ICAR publication provides age banded tables. Young, Keith and Bond examined age and sex invariance of the ICAR precisely because it had not been established, which is the responsible order of operations and also an admission that the groundwork was not yet done.

The reference distribution moves. This one gets its own section below, because the evidence for it comes from the ICAR's own data and is more interesting than a caveat.

What you can legitimately say about an ICAR scoreYou can say how many items you answered correctly. You can compare your score to the mean of a specific published sample if you name that sample and its selection method. You can compare two groups within one study who took the identical form. You cannot state an IQ, a percentile, or a rarity, and neither can anyone else, however confidently a forum post asserts otherwise. If a percentile is what you need, an instrument that publishes an actual percentile table and a rarity figure is the tool for the job, not a research item bank.

Notice what is absent from this list. Nobody is arguing that the ICAR measures poorly. The measurement evidence is presented two sections down and it is good. The missing thing is a norm, and a norm is a separate deliverable that costs money the project was never funded to spend.

7 What It Would Take to Publish an ICAR Norm

Norming is not a statistical afterthought applied to data you already have. It is a data collection project with its own design, its own budget and its own expiration date. Setting out the requirements makes it obvious why a grant funded public domain project has not produced one.

  • A named target population and a sampling plan. Adults aged 16 to 90 in a specified country, sampled against known population characteristics, with documented recruitment and documented refusals. Convenience samples cannot be repaired after the fact by weighting alone.
  • A frozen form. One fixed item set, one fixed order, one fixed administration mode. The moment researchers are free to pick items, the norm no longer describes what anyone actually took.
  • Stratification and documented exclusions. Age bands, education bands, and a published statement of who was excluded and why. Exclusion rules move norms substantially, which is why they are argued about in every test manual.
  • Measurement invariance evidence. Proof that the score means the same thing across the groups being compared, before percentiles are handed to people in those groups. This is the technical core of the question of whether a test is biased, where differential item functioning and invariance testing do the work that intuition cannot.
  • A published standard error and confidence interval. A percentile without a confidence band invites a precision that the instrument does not have.
  • A date and a re-norming plan. Norms decay. A table with no year on it is a table nobody can evaluate.

Commercial publishers fund this from royalty income. That is the actual trade the 2014 ICAR paper describes: royalty streams have funded the development, distribution and maintenance of commercial measures for decades, and a public domain resource by definition forgoes that stream. The ICAR chose free items and open data. The cost of that choice, paid by anyone holding a raw score, is that no norm ever arrived.

This is also why "just build a norm from the SAPA data" does not work. Those data were collected by matrix sampling from self selected volunteers, so they violate the first two requirements above before you begin. You would be publishing the percentile rank of a person among internet volunteers who took an overlapping subset of items, dressed up as a population statement. The intellectual honesty of the ICAR team is visible in the fact that they never did it.

8 How Strong the ICAR Evidence Actually Is

Strip away the norm question and the ICAR's construct evidence compares well with instruments that charge money. This section exists so that the article's central criticism is not mistaken for a general dismissal.

The single most important validation is Young and Keith (2020), published in the Journal of Psychoeducational Assessment and catalogued at its ERIC record. They administered the ICAR 16 and the Wechsler Adult Intelligence Scale, Fourth Edition, to a convenience sample of 97 students. The correlation between observed ICAR 16 total and WAIS-IV Full Scale IQ was .81. Between the confirmatory factor analytic general factors of the two instruments, the correlation was .94. Their conclusion was that the ICAR 16 is a valid brief measure of nonverbal intelligence, with the explicit caveat that replication in larger samples is needed. That wording is worth holding on to, because a nonverbal estimate is a narrower claim than a general one, and what a nonverbal IQ test can and cannot cover includes the reminder that language still enters through instructions and conventions. Their factor results also complicate a common assumption: the letter and number series task loaded on fluid reasoning, while matrix reasoning, verbal reasoning and three dimensional rotation loaded on visual spatial ability. Anyone treating ICAR matrix items as a pure fluid measure should read that finding before they do.

ComparisonCorrelationSampleSource
ICAR 16 total with WAIS-IV Full Scale IQ.81 observed, .94 between latent general factors97 studentsYoung and Keith, 2020
ICAR 16 with Shipley-2 composites.41 uncorrected, .82 and .81 after range and reliability correction137 university studentsCondon and Revelle, 2014
ICAR 60 with self reported combined SAT.54 uncorrected34,229 online participants aged 18 to 22Condon and Revelle, 2014
ICAR 60 with self reported ACT.49 uncorrected12,254 reporting an ACT scoreCondon and Revelle, 2014
ICAR Puzzle Completion with Raven's Progressive Matrices.40100 US adults aged 18 to 82Young and colleagues, 2025
ICAR Block Rotation with Woodcock Johnson Visualization.51100 US adults aged 18 to 82Young and colleagues, 2025

The two achievement rows carry a caution. Those SAT and ACT figures were self reported rather than verified, and the relationship between an IQ score and an SAT result is an estimated correspondence between two instruments, not a fixed exchange rate. A correlation near .5 with a self reported entrance score is respectable evidence about a construct and useless as a scoring rule.

The most recent entry deserves a note. Young, Kim, McKee, Doyle, Novack, Revelle, Gershon and Dworak (2025), in the Journal of Intelligence, implemented two ICAR measures inside the Mobile Toolbox platform: a 30 item fixed form built from the progressive matrices content, renamed Puzzle Completion, and a computer adaptive rotation test capped at 25 items, renamed Block Rotation. In 100 US adults aged 18 to 82, empirical reliability was .88 and .81 respectively, with alpha .90 and .93 and omega total .90 and .89, and practice effects on retest were negligible. And still, scoring was reported as item response theory theta, with no standardized transformation and no normative data. A 2025 paper by the ICAR team's own collaborators, on a modernized adaptive implementation, declines to publish a norm. That is the clearest evidence available that the absence is a choice about scope rather than an oversight. For the broader question of what separates a well measured test from an accurate score, see what makes an IQ test accurate.

9 The Exposure Problem a Public Item Bank Has

A proprietary item bank is protected by a paywall and a credible legal threat. A public domain bank has neither, and its authors have been unusually candid about what that costs.

The 2014 paper treated exposure as the central theoretical objection to the whole project. Condon and Revelle wrote that the lack of copyright protection poses a theoretical threat to test validity, then argued the magnitude of that threat was unknown and could be offset. Their third development goal was to avoid item content that could be readily referenced elsewhere. Their proposed mitigation was counterintuitive but coherent: make sample items easily available to the general public, on the theory that visible sample items discourage wholesale distribution of the full set and reduce the advantage held by whoever has seen the bank. Their empirical claim was that after administration to 97,000 participants, the public domain status of the items did not compromise their validity.

Six years later the team was blunter. In the Personality and Individual Differences overview, Dworak and colleagues state that despite their best efforts to keep the items secure, copies have appeared on the internet. They also record a design decision that follows directly from this: vocabulary items are intentionally missing from the pool, because in an unsupervised web session it is trivial to solve a vocabulary problem with a search engine. A test author leaving out an entire highly g loaded content domain because he cannot defend it against a browser tab is telling you something real about unsupervised testing generally.

The structural fix they identify is automatic item generation: algorithms that produce fresh items with known difficulty parameters, so that the universe of possible items is effectively unbounded and no leaked set matters. Progress exists for arithmetic reasoning, figural analogies and perceptual mazes, and the ICAR group continues the work. It is not finished, and the four classic item types are not generated on the fly.

When exposure matters and when it does notFor a correlational study with cooperative volunteers who gain nothing by cheating, exposure is close to irrelevant. It becomes decisive the moment a score carries a consequence: selection, admission, a claim of membership, or a number a person will repeat about themselves. This is why Raven's 2 controls item exposure through its digital forms and why any comparison of online tests should ask what the item bank is protected by, not just how good the items are. It is also the failure mode behind most high range tests, whose banks circulate freely and whose numbers inflate accordingly.

There is a second layer. The scored responses of 1,525 participants on the ICAR 16, along with the item identifiers, ship inside an R package that anyone can install. That is superb open science and it is also a permanent public record of which 16 items constitute the standard short form. Both things are true, and the ICAR team would say both out loud.

10 What a Moving Reference Sample Does to an Old Score Distribution

Even if someone built a percentile table from the 2014 distribution today, it would already be wrong, and the evidence comes from the ICAR's own data.

Dworak, Revelle and Condon published Looking for Flynn effects in a recent online U.S. adult sample in Intelligence in 2023. They analyzed thirteen years of cross sectional SAPA data from 394,378 United States adults collected between 2006 and 2018, using an overlapping set of 35 ICAR items collected across the whole period and 60 items collected from 2011 onward. The question was whether average cognitive ability scores in this sample moved over time.

They did move, and not in one direction. Composite scores from the 35 item set, and the domain scores for matrix reasoning and for letter and number series, showed a pattern consistent with a reversed Flynn effect from 2006 to 2018 when stratified by age, by education or by gender. Composite scores from the 60 item set showed the same reversal from 2011 to 2018. Verbal reasoning slopes did not meet or exceed the authors' annual threshold of 0.02 standard deviations, so verbal was effectively flat. Three dimensional rotation went the opposite way, showing a rising pattern, with the largest slopes in age stratified regressions.

Read that carefully, because it does two things at once. It supplies real evidence about population level score change, which is the topic the Flynn effect page covers in general. And it demolishes the idea of a single fixed ICAR conversion table, because different item types drifted in different directions within the same instrument over the same period. A composite conversion built in 2014 would now overstate rotation performance and understate matrix performance relative to a contemporary comparison group, and the size of the error would depend on which form a person took.

The authors themselves are careful about what their sample can support. They note that participation was voluntary and the sample non representative of the target population, and they contrast this directly with the systematic norming data collected by proprietary licensed measures using probability sampling or population based conscript data. They even raise the possibility that the composition of SAPA volunteers changed as the project was discussed in social media and online articles, so that later cohorts may be more ordinary than earlier ones. That is a sober caveat, and it is fatal to using their series as a norm while being perfectly compatible with the paper's actual argument about trends.

11 When the ICAR Is the Right Tool and When It Is Not

The ICAR is excellent at the job it was built for and unusable for several jobs people try to give it. The dividing line is simple: the ICAR compares people to each other inside your study, and it cannot compare a person to a population outside it.

What you are trying to doIs the ICAR appropriateWhy
Control for cognitive ability as a covariate in a correlational studyYesRelative standing within your sample is all a covariate needs, and the reliability evidence supports it.
Compare two experimental conditions on reasoning performanceYesBoth groups take the identical form, so the absence of a norm is irrelevant to the contrast.
Run a large online study on a budget with no licensingYesThis is the exact use case the 2014 paper was written to enable.
Tell an individual participant where they standNoNo reference population, no fixed form, no age stratification, no published percentile.
Screen for giftedness or cognitive impairmentNoScreening needs cut scores against a defined population, which is why adult giftedness assessment starts from a normed instrument.
Support hiring, admissions or accommodation decisionsNoUnsupervised administration, an exposed bank and no norm rule this out, as they do for any unsupervised test used for pre employment cognitive screening.
Publish a percentile or an IQ figure in a reportNoThere is no defensible transformation from a raw ICAR count to either metric.
Give participants a score they can take awayNoThe project deliberately produces no interpretive output, and improvising one creates a claim you cannot support.

The last row is where most real problems start. A researcher runs a study, participants ask what their score meant, and the temptation is to hand back something. Handing back a raw count with an explanation of its limits is fine. Handing back a percentile invented from the study sample is not, because the study sample is not a population, and the number will outlive the caveat that accompanied it. If participant feedback is part of your design, that requirement should drive instrument choice from the beginning rather than being patched on at the end. Choosing the instrument also means choosing a product category before comparing vendors, since only one of the four categories that share the name cognitive assessment platform returns a composite on a population scale at all, which is the sorting the survey of administration platforms by what they output does before any feature grid. The distinction between administration contexts is covered on the professional versus online administration page, and the question of where a normed adult score can actually be obtained on the page on where to take an IQ test.

12 How ACIS Professional Handles the Same Job

ACIS solves a different half of the problem than the ICAR does, and the honest framing is a trade rather than a ranking. The ICAR gives away its items and publishes its data. ACIS keeps its item banks closed and publishes a normed score with a stated error term. Neither choice is free.

Structurally, ACIS is a 20 subtest battery across six primary cognitive domains, reporting subtest scaled scores on a mean of 10 with a standard deviation of 3, and composites on a mean of 100 with a standard deviation of 15. Its documented figures come from a technical analysis set of 2,750 complete records, within an adult reference frame of 3,243 English speaking records covering ages 16 to 90. The Full Scale IQ composite has an omega of .9886 and a g loading of .958, with a standard error of measurement of about 1.60 IQ points. The higher order g confirmatory model fit indices are CFI .9761, TLI .9726, RMSEA .0406, SRMR .0217, with a chi-square of 916.703 on 166 degrees of freedom. All of these are set out in the ACIS technical manual.

The comparison is not like for likeA 16 item short form and a 20 subtest composite are different kinds of object, and a smaller standard error on the longer instrument is arithmetic rather than an argument. The relevant ACIS limitations must be stated alongside those figures: the normative sample is self selected rather than census based, administration is unsupervised, the adult reference frame is a modelled frame rather than a census sample, and ACIS is not a clinical or diagnostic instrument. It is not appropriate for diagnosis, hiring decisions, accommodations, or high IQ society admission.

For a research team, the practical difference is what leaves the session. An ICAR administration produces a raw count you must interpret yourself with no defensible population statement. An ACIS Professional administration produces a scored profile with subtest scaled scores, six index composites and a Full Scale IQ carrying a published standard error, which means a participant report and a dataset column can both exist without anyone inventing a percentile.

ACIS Professional is a workspace rather than a checkout. A free account includes three Quick administrations, after which administrations are pay as you go at the same prices as the consumer tiers. Participants receive a private link, no participant personal data is collected, scores are verified rather than self reported, and results export to CSV. It is intended for research, screening and educational use, and explicitly not for clinical decisions. The full description is on the ACIS for research page, and the same pitch appears at the administer section of the home page. If your question is whether a scored composite is worth the closed item bank, the answer depends entirely on whether anyone downstream needs to read a number and know what it means, which is the question the Full Scale IQ page addresses directly.

13 Reading an ICAR Score Responsibly

If you are holding an ICAR score right now, the responsible reading is narrow and it is still useful. Write down the exact form you took, the number of items, which item types were included, whether it was proctored, and what you scored. That set of facts is your result. Anything beyond it is inference you cannot support.

In a manuscript, the reporting standard is not difficult. Name the item types and the exact item count. State that administration was unsupervised and untimed if it was. Report the internal consistency computed in your own sample rather than quoting Condon and Revelle's coefficients as though reliability were a property the instrument carries between studies. Treat the score as an interval level covariate or a group comparison, and say explicitly that no normative interpretation is available. Reviewers accept this readily, because it is what the instrument's own authors do.

What not to write is equally short. Do not report an IQ. Do not report a percentile. Do not describe a participant as above average without naming the average you mean and where it came from. Do not compare an ICAR score to a WAIS or Stanford-Binet figure as though they sit on a shared scale, a point that applies equally when reading about the WAIS-5 and its index structure or about the way the Stanford-Binet 5 builds its composites. Do not pool ICAR scores across studies that used different item subsets. And do not let a raw count acquire a percentile in the retelling, which is the specific failure this whole page is about.

These are not house rules. The Standards for Educational and Psychological Testing (2014), published jointly by the American Educational Research Association, the American Psychological Association and the National Council on Measurement in Education, make the requirement explicit: norms must be described with reference to the population sampled, the sampling method and the date of collection, and a validity argument supports one specified interpretation for one specified use rather than the instrument in the abstract. APA testing standards apply the same logic to who may interpret a score and under what conditions. Measured against those requirements, the ICAR has a strong reliability and validity record for research interpretation, and no normative claim available for individual interpretation. That is not a failing grade. It is an accurate description of an instrument whose authors documented their own limits more carefully than most publishers do, and reading it any other way is the reader's error rather than theirs.

14 Frequently Asked Questions

What is the ICAR test?

It is the International Cognitive Ability Resource, a public domain pool of cognitive ability items maintained by an academic consortium and introduced by Condon and Revelle in 2014. Researchers assemble their own forms from the pool rather than buying a fixed product.

Is the ICAR a real IQ test?

It is a real measure of cognitive ability with published reliability and validity evidence, but it is not an IQ test in the sense of producing a standardized IQ score, because it has no normative table.

How many items are in the ICAR 60?

Sixty, split unevenly: 24 three dimensional rotation, 16 verbal reasoning, 11 matrix reasoning and 9 letter and number series items.

What is the ICAR 16?

It is the ICAR Sample Test, a short form containing four items from each of the four original types. It is the version most commonly administered in published research.

Can I convert an ICAR 60 score to an IQ?

No. There is no published conversion, and building one would require a stratified reference sample and a frozen form that the project has never had.

Are there ICAR 60 norms anywhere?

Not in the peer reviewed literature or on the project site. The large SAPA sample supports item level analysis, not population percentiles.

Is the ICAR 60 timed?

No. The ICAR measures are power tests and timing plays no part in administration or scoring, although researchers sometimes impose a practical upper limit.

How long does the ICAR 16 take?

The ICAR team suggests capping it at 16 minutes for adults aged 18 to 25, and notes that most participants finish in roughly half that time.

Who created the ICAR?

David Condon and William Revelle at Northwestern University introduced it, with subsequent development by an international consortium funded through the NSF in the United States, the DFG in Germany and the ESRC in the United Kingdom.

Is the ICAR free to use?

Yes for non commercial research purposes. The content repository requires user registration rather than offering an anonymous bulk download.

How reliable is the ICAR 16?

Condon and Revelle report an alpha of .81 and an omega total of .83 in their founding sample. That is adequate for a 16 item form and modest by the standards of a full battery.

Why is ICAR matrix reasoning the weakest subscale?

Its 11 items produced an alpha of .68, and the authors reported that they were not uniformly measuring a single latent construct. The psychTools documentation adds that two items lack monotonically increasing trace lines.

How well does the ICAR correlate with the WAIS?

In Young and Keith's 2020 study of 97 students, the observed ICAR 16 total correlated .81 with WAIS-IV Full Scale IQ, and the latent general factors of the two instruments correlated .94.

Does the ICAR measure fluid reasoning?

Partly. Young and Keith found the letter and number series task loading on fluid reasoning while matrix, verbal and rotation tasks loaded on visual spatial ability, which is not what most users assume.

Can I use the ICAR to screen job applicants?

No. Unsupervised administration, an openly available item bank and the absence of norms make it unsuitable for any decision that affects a person's opportunities.

Where can I find a sample ICAR test?

The project site offers a sample test behind registration. Condon and Revelle argued that making sample items publicly visible actually discourages wholesale distribution of the full bank.

Why does the ICAR have no vocabulary items?

They were deliberately excluded. The ICAR team notes that in an unsupervised web session a vocabulary item is trivially solved with a search engine.

Have ICAR items leaked online?

Yes. In their 2020 overview the ICAR team states that despite their efforts to keep items secure, copies have appeared on the internet, and points to automatic item generation as the long term fix.

Did ICAR scores change over time?

In 394,378 US adults sampled between 2006 and 2018, Dworak, Revelle and Condon found declining composite, matrix and letter series scores, flat verbal scores, and rising three dimensional rotation scores.

Is the ICAR better than a paid online IQ test?

It is better documented than most, and it delivers less to an individual, because it deliberately produces no interpretive score. Which matters more depends entirely on whether anyone needs to read the result.

What should I do if I need an actual normed score?

Use an instrument that publishes a reference frame, a standard error and a dated normative table, and check that the intended use is one the publisher supports. A research item bank is the wrong tool for that request.

Take the assessment

You get a profile, not a number

ACIS measures six CHC domains across 20 subtests and reports each one with its own normed score and confidence interval, so you can see where you are strong and where you are not.

Free trial, no card required. Full report from $15.