Three Item Measure

The Cognitive Reflection Test and its three answers

Three questions, no time limit, arithmetic a sixth grader can do. The Cognitive Reflection Test is the shortest instrument in wide use in behavioral science, it is routinely described online as the world's shortest IQ test, and it is not an IQ test. This page gives the items, works each answer through, and reports what the published data say the score is worth.

A woman at a desk with a pencil held to her temple, looking down at a page of notes beside a blue mug and a stack of books.
Frederick administered the three items to 3,428 people between 2003 and 2005, and a third of them scored zero.

0 Quick Answer

The Cognitive Reflection Test is three problems long, it was built to measure whether you check your first answer rather than how much you can compute, and it has never had norms. Shane Frederick published it in the Journal of Economic Perspectives in 2005, in volume 19, issue 4, pages 25 to 42. The three items are the bat and ball problem, the machines and widgets problem, and the lily pads problem, and all three are printed in Figure 1 of that paper. They are reproduced and solved in the next section of this page.

The test went to 3,428 respondents in 35 separate studies over a 26 month period beginning in January 2003, most of them undergraduates paid eight dollars to complete a 45 minute questionnaire. The overall mean was 1.24 correct out of three. Thirty-three percent of respondents got none of the three right, 28 percent got one, 23 percent got two, and 17 percent got all three. Those are convenience samples of paid undergraduates and passers by, not a norm group, so they describe who took the test rather than the population.

The test correlates with measured ability, but moderately. In Frederick's Table 4, the CRT correlated .44 with self reported SAT total scores across 434 respondents, .46 with ACT scores across 667, .43 with the Wonderlic Personnel Test across 921, and .22 with the Need For Cognition scale across 944. That is enough shared variance to show the tests draw on common factors and far too little to treat three items as a substitute for a battery. What that gap means in practice is set out on the page on what an intelligence test measures.

1.24 of 3

Overall mean CRT score across 3,428 respondents in 35 studies, Frederick 2005, Table 1.

33 percent

Share of that pooled sample who answered none of the three items correctly.

.43 to .46

Correlations with the Wonderlic, the SAT and the ACT reported in Frederick's Table 4, which is roughly 18 to 21 percent shared variance.

1 The Three Items and How Each One Works

Every item is designed so that a wrong answer arrives before you have decided to answer, and so that the correct answer needs no mathematics beyond middle school. All three appear in Figure 1 of Frederick's 2005 paper, which is the reason they can be printed here. What follows is each item as published, the answer that arrives first, the answer that is correct, and the step that separates them.

Item one, the bat and ball. A bat and a ball cost $1.10 in total. The bat costs $1.00 more than the ball. How much does the ball cost? The answer that springs to mind is 10 cents. It is wrong, and the check takes one subtraction: if the ball were 10 cents, the bat at a dollar more would be $1.10, the pair would come to $1.20, and the stated total is $1.10. Solve it properly and the ball is x, the bat is x plus 1.00, so 2x plus 1.00 equals 1.10, which gives x equal to 0.05. The ball costs 5 cents, the bat costs $1.05, the difference is exactly $1.00, and the total is exactly $1.10. Frederick noted that nearly everyone who does not answer 10 cents answers 5 cents, which means catching the error is almost the same thing as solving the problem.

Item two, the machines. If it takes 5 machines 5 minutes to make 5 widgets, how long would it take 100 machines to make 100 widgets? The answer that arrives is 100 minutes, produced by scaling every number in the sentence at once. The correct answer is 5 minutes. The quantity that stays fixed is the rate of a single machine, not the total output. Five machines making five widgets in five minutes means one machine makes one widget in five minutes. A hundred machines running in parallel therefore make a hundred widgets in the same five minutes, because each machine is still finishing its own widget in five minutes. Doubling the number of workers does not double the time each worker needs.

Item three, the lily pads. In a lake, there is a patch of lily pads. Every day, the patch doubles in size. If it takes 48 days for the patch to cover the entire lake, how long would it take for the patch to cover half of the lake? The answer that arrives is 24 days, produced by halving the stated number. The correct answer is 47 days. Doubling means each day's area is twice the previous day's, so the day before full coverage is by definition the day of half coverage. Working forward from day one requires you to track 48 doublings. Working backward from day 48 requires one step. On day 24 the patch is not half the lake, it is roughly one sixteen millionth of it, because it still has 24 doublings to go.

Why printing the answers does not spoil anythingThe items have been public since 2005 and reproduced in thousands of places since. The familiarity research summarized further down this page found that between 44.9 percent and 65.6 percent of participants in recent samples already knew them. A page that withheld the answers would not protect the instrument, it would only make the reader go find them somewhere less careful. The relevant question is what a score means once everybody has seen the items, and that question has an answer.

Notice what the three problems have in common. Each has an incorrect response that is generated by an obvious surface operation: subtract a dollar, scale everything, halve the total. Each has a correct response that requires one act of checking. And each is trivially easy to explain in a sentence once the trick is named. Frederick built them that way on purpose, because a problem that is genuinely hard to compute would measure computation instead. The formats used by a scored ability battery, where difficulty is the point rather than the trap, are described on the page on IQ test question types.

2 What Frederick Said He Was Measuring

Frederick defined cognitive reflection as the ability or disposition to resist reporting the response that first comes to mind, and the whole design follows from that definition. He contrasted the CRT items with a problem that offers no intuitive answer at all: finding the square root of 19,163 to two decimal places without a calculator. That task requires effort, motivation, and a learned procedure, but no number springs to mind, so there is nothing to override. It measures capacity with no reflective component. The bat and ball measures the opposite thing, because the computation is trivial and the only difficulty is the answer that arrives uninvited.

He gave four pieces of evidence in the 2005 paper that the wrong answers really are intuitive rather than merely wrong.

  • Among all the wrong answers people could give, the three posited intuitive answers, 10, 100 and 24, dominated the alternatives.
  • Even among people who answered correctly, the wrong answer often came first. Frederick reported that 10 cents was often crossed out next to 5 cents on the answer sheets, but never the other way around.
  • People who missed the items thought they were easier than people who solved them. Respondents who answered 10 cents estimated that 92 percent of others would solve the bat and ball problem. Respondents who answered 5 cents estimated 62 percent. Both figures were substantial overestimates.
  • Respondents did much better on an analogous problem built to invite computation. Frederick's control item asks about a banana and a bagel costing 37 cents together, with the banana 13 cents more than the bagel. The numbers are awkward enough that nothing springs to mind, so people calculate, and they get it right far more often than they get the bat and ball right.

That fourth point is the sharpest of the four, because it isolates the variable. The banana and bagel problem is arithmetically harder than the bat and ball problem and is answered correctly more often. Whatever the CRT is measuring, it is not the difficulty of the sum. It is whether the person notices that a sum is required at all.

This is also the reason the instrument sits awkwardly next to intelligence testing. A battery presents a problem, signals unambiguously that a careful answer is expected, applies a time limit, and measures how well the machinery runs. The CRT hides that signal inside a problem that looks too easy to be worth checking. The separate question of whether thinking well and scoring well come apart in the same person is argued from the rationality literature on the critical thinking versus intelligence page, and this page stays with the instrument itself.

3 What People Actually Score, by Sample

Table 1 of Frederick's paper is the closest thing the CRT has to published norms, and it is a table of eleven convenience samples rather than a reference distribution. The spread across those samples is large. The highest mean, 2.18 out of three, came from 61 respondents at the Massachusetts Institute of Technology, of whom 48 percent answered all three items correctly. The lowest, 0.57, came from 138 respondents at the University of Toledo, of whom 64 percent answered none correctly. Everything about the ordering is consistent with selectivity of admission, which is a point about the samples rather than about the instrument.

Location where data were collectedMean CRT scoreScored 0Scored 3N
Massachusetts Institute of Technology2.187%48%61
Princeton University1.6318%26%121
Boston fireworks display1.5324%26%195
Carnegie Mellon University1.5125%25%746
Harvard University1.4320%20%51
University of Michigan, Ann Arbor1.1831%14%1,267
Web based studies1.1039%13%525
Bowling Green University0.8750%12%52
University of Michigan, Dearborn0.8351%6%154
Michigan State University0.7949%6%118
University of Toledo0.5764%5%138
Overall1.2433%17%3,428

Three footnotes in the original change how several rows should be read. The Boston fireworks respondents were people waiting on the grass for a July 4th display, they received a frozen ice cream bar rather than the usual eight dollars, their ages varied more widely than in the student samples, and some of them completed the survey in small groups where other people had already done it. The Harvard sample of 51 was entirely female, unlike every other location. The web based studies drew unpaid participants whose email addresses came from online retailers, entered into a prize lottery instead of being paid.

The reason to publish the table anyway is that people search for a CRT average and find numbers with no provenance attached. These have provenance. What they do not have is a normative frame: no age adjustment, no representative sampling, no standard score, and no way to convert a raw count into a percentile of anything except these eleven groups. If the comparison you want is between institutions rather than between individuals, the evidence on that question is assembled on the average IQ by university page, and the machinery that turns a raw score into a population percentile is described on the norming page.

4 How the CRT Relates to Measured Ability

The correlations are real, consistent, and moderate, and the last word is the one that matters. Of the 3,428 respondents who took the CRT, many completed a second cognitive measure as part of the same questionnaire: 921 took the Wonderlic Personnel Test, a 12 minute, 50 item test used by employers including the National Football League, 944 completed an 18 item Need For Cognition scale, and several hundred reported SAT or ACT scores. Frederick's Table 4 reports the resulting correlation matrix along with the sample size behind each cell.

.44 with SAT

CRT against self reported SAT total, N = 434. Against SAT math it was .46, against SAT verbal .24.

.43 with Wonderlic

CRT against the Wonderlic Personnel Test, N = 921. The Wonderlic itself correlated .49 with SAT total.

.22 with Need For Cognition

CRT against the self report thinking disposition scale, N = 944, the weakest association in the matrix.

Read the SAT breakdown carefully, because it is the most informative line in the table. The CRT correlated .46 with SAT math and .24 with SAT verbal. A measure of reflection that had nothing to do with numbers would not show that split. Frederick said so himself: performance on the CRT is surely aided by reading comprehension and mathematical skill, which is exactly what the ACT and SAT are built to measure. Later work made the same point from the other direction, with Campitelli and Gerrans demonstrating through mathematical modeling in 2014 that the CRT is not simply a numeracy test, which would not have needed demonstrating if numeracy had not been a live explanation.

A correlation near .45 means roughly 20 percent shared variance and about 80 percent that is not shared. In practical terms, knowing a person's CRT score narrows the plausible range of their ability score a little and leaves most of it open. That is the general shape of the relationship between any short measure and a broad one, and it is why composite scores are built from many subtests rather than from the best single task. The statistical core behind that argument is set out on the g factor page, and if you want to see what a given ability score corresponds to in population terms you can convert a score to a percentile directly.

5 Why It Is Not an IQ Test

Four separate facts each disqualify the CRT as an intelligence test, and they compound. The phrase world's shortest IQ test attached itself to the instrument through repetition rather than through any claim its author made. Frederick called it a simple measure of one type of cognitive ability. One type is doing a great deal of work in that sentence.

It has no normative frame. An IQ score is a rank against a reference sample of people the same age, expressed on a scale with a fixed mean and a fixed standard deviation. The CRT produces a count from zero to three. There is no reference sample, no age adjustment, and no conversion table, so there is nothing to be 100 or 130 relative to. Why the 15 point scale exists and what it buys is explained on the standard deviation page.

Three dichotomous items cannot carry an individual score. Frederick did not report a reliability coefficient in the original paper, and as Primi and colleagues noted in the Journal of Behavioral Decision Making in 2016, most researchers who adopted the scale followed the same practice. The figures that have been published sit in a range that would be unacceptable for individual decisions: Liberali and colleagues reported Cronbach's alpha of .74 in one study and .64 in another, Weller and colleagues .60, Campitelli and Gerrans .66, and Morsanyi and colleagues .57 and .68. Thomson and Oppenheimer reported .624 for the original three items in Judgment and Decision Making in 2016. Why that number governs what a score can be used for is covered on the reliability and validity page.

The correlations with ability are moderate, not high. The .43 to .46 range in the previous section is the ceiling of the published evidence, not a conservative estimate. Primi and colleagues found correlations of .32 between the original CRT and Set I of Raven's Advanced Progressive Matrices in a subsample of 201 participants, which is lower still.

It measures a disposition to check, not a capacity to solve. This is the point most often lost. Every CRT item is arithmetically trivial. Nobody fails the bat and ball problem because they cannot subtract. They fail it because they answered before they decided to answer. A test on which the computation is deliberately made easy cannot rank people on how much computation they can do.

What a three item score can and cannot doAggregated over hundreds of people, a three item measure can detect group differences and predict group level outcomes, which is why the CRT has been productive in research. For one person on one occasion, three binary items produce a score with a wide error band, no reference distribution, and no interpretation beyond how many of three specific problems that person got right. Both statements are true at once, and confusing the first for the second is how a research instrument became an internet IQ test.

6 The Items Are Now the Most Contaminated in Psychology

A trick problem stops being a trick problem once you have seen the trick, and by the mid 2010s a large share of every research sample had. Four studies measured this directly by simply asking participants whether they had encountered the items before, and their estimates converge.

51.4 percent

Share of 142 online volunteers who had previously seen at least one CRT problem, Haigh 2016, Advances in Cognitive Psychology.

44.9 percent

Share of 2,137 participants reporting prior encounter with the CRT or similar tasks, Stieger and Reips 2016, PeerJ.

65.6 percent

Share of Mechanical Turk participants who had seen the bat and ball problem before, Woike 2019, Frontiers in Psychology.

Matthew Haigh reported in Advances in Cognitive Psychology, volume 12, issue 3, pages 145 to 149, that his 142 volunteers split almost evenly on prior exposure, and that the split mattered enormously. Participants who had seen at least one item scored a mean of 2.36 out of three with a standard deviation of 0.96. Participants who had not scored 1.48 with a standard deviation of 1.21. The difference tested at t(140) equal to 4.802 with p below .001 and a Cohen's d of 0.81, which is a large effect by any convention.

Stefan Stieger and Ulf Dietrich Reips reported the same pattern at scale in PeerJ in 2016, with a final sample of 2,137. Experienced participants averaged 1.65 against 1.21 for naive participants, t(2135) equal to 9.38, p below .001, d equal to 0.41. Their conclusion was that the three item form is limited both by familiarity and by range restriction, and they recommended longer forms.

Thomson and Oppenheimer found the contamination concentrated exactly where research runs. In their 2016 paper, 62.6 percent of one Mechanical Turk sample reported prior exposure to the original CRT, and among workers who had qualified for Masters status on the platform, meaning experienced high volume respondents, the figure was 94.0 percent. Jan Woike's 2019 study in Frontiers in Psychology, covering 3,660 participants across three studies, found that 65.6 percent of Mechanical Turk respondents had seen the bat and ball problem, and that 93.7 percent of those had first encountered it on the platform itself.

This is a structural problem, not a fixable one. Three items with memorable answers cannot be un published. Once a person knows that the ball costs 5 cents, their answer records a memory rather than a reflection, and no scoring rule can tell the two apart. It is the same failure mode that makes any widely circulated item set unusable, which is one of the reasons a scored battery keeps its content off the open web and one of the differences catalogued on the free versus validated comparison.

7 Whether Familiarity Actually Breaks the Measure

Exposure raises the score without necessarily destroying the ordering, and the distinction decides who should worry. Michal Bialek and Gordon Pennycook tested this directly and published the result in Behavior Research Methods, volume 50, pages 1953 to 1959. Across six studies with roughly 2,500 participants and 17 outcome variables including religious belief, receptivity to pseudo profound statements, smartphone usage, numeracy, and susceptibility to heuristics and biases, they compared how strongly the CRT predicted each outcome among participants with and without prior exposure.

They replicated the score inflation. About a third of their pooled sample reported prior exposure, and those participants scored higher, with a Cohen's d of 0.57, in line with the 0.48 and 0.41 reported by Haigh and by Stieger and Reips. Then they compared the predictive correlations. Of 23 comparisons tested with Fisher's z, only three differed significantly between the experienced and inexperienced groups, and in all three cases the CRT was more predictive among the participants who had seen it before, not less.

Their proposed explanation is worth stating because it is testable and slightly unsettling. Being handed an apparently trivial problem for the second time is itself information about the problem's difficulty. Reflective people notice that and slow down. Less reflective people do not, and answer 10 cents again. On that account, repeated exposure does not erase the individual difference the test is trying to catch, it re expresses it.

The honest reading, which cuts both waysFor a researcher pooling hundreds of participants, the published evidence says the CRT survives contamination well enough to keep using, with parallel forms and an exposure question recommended as controls. For one individual who has read about the bat and ball problem and then takes the test, the score is uninterpretable, because there is no way to separate what they worked out from what they remembered. A group statistic that is robust and an individual score that is meaningless are perfectly compatible, and most of the confusion around this instrument lives in the gap between them.

The same asymmetry appears elsewhere in testing and is worth recognizing as a pattern. Practice effects, coaching, and item exposure all tend to shift a distribution upward while leaving much of the rank order intact, which is why a single person's improvement after preparation is weak evidence about their ability and a population shift is strong evidence about something. Several related claims that survived on repetition rather than evidence are catalogued on the myths page.

8 The Replacement Items and Longer Forms

Three research groups responded to the contamination by building new items rather than defending the old ones, and their solutions differ in an instructive way. One removed the arithmetic, one added parallel items in the same style, and one extended the difficulty range downward so the scale would work outside selective universities.

Version and sourceItemsReported reliabilityWhat problem it was built to solve
Original CRT, Frederick, Journal of Economic Perspectives, 20053Not reported by the author; .624 in Thomson and Oppenheimer 2016Nothing. It was the first version.
CRT-2, Thomson and Oppenheimer, Judgment and Decision Making, 20164alpha .511 alone, .705 combined with the originalPrior exposure, and the heavy numerical loading of the original items
Expanded CRT, Toplak, West and Stanovich, Thinking and Reasoning, 20144 new, 7 combined.72 for the seven item compositePrior exposure, and the coarseness of a three point scale
CRT-Long, Primi and colleagues, Journal of Behavioral Decision Making, 20166, being 3 original plus 3 newalpha .76 in the main sample, .79 in an adolescent sampleFloor effects outside highly educated adult samples

Keith Thomson and Daniel Oppenheimer built CRT-2 around problems whose intuitive pull comes from language and framing rather than from arithmetic, so that mathematical skill would contribute less. Their items ask about the finishing position in a race, a count of surviving sheep, the name of a fourth daughter, and the volume of a hole, and each has a fast wrong answer and a slower right one. Across a Mechanical Turk sample of 200 and a UCLA undergraduate sample of 143, CRT-2 correlated .511 with the original CRT, which rose to .905 after correcting for the unreliability of both short scales. That disattenuated figure is the important one: it says the two sets of items are measuring close to the same thing, despite sharing no content.

Maggie Toplak, Richard West and Keith Stanovich took the other approach in Thinking and Reasoning in 2014, volume 20, pages 147 to 168, adding four items in the original style and reporting that the seven item combination reached a reliability of .72. They were explicit about the motivation, writing that the original three items were becoming known to potential participants.

Caterina Primi, Kinga Morsanyi, Francesca Chiesi, Maria Anna Donati and Jayne Hamilton attacked a different limitation. Using a two parameter item response theory model on 438 university students in Florence and Belfast and then on a second sample of 988 students, they showed that the original items are difficult enough to produce a floor: 30 percent of their sample scored zero on the three item scale, against 9 percent on the six item CRT-Long. Their published report gives the validity evidence as well, with the CRT-Long correlating .39 with Raven's Advanced Progressive Matrices Set I against .32 for the original, and .44 with an objective numeracy scale against .41.

The pattern across all three is the same admission. A three item test is too short to rank individuals, too easy to remember to stay uncontaminated, and too hard to discriminate below the top of the range. Building items that discriminate at a specified difficulty is a slow, iterative business, and what it looks like when the target is the upper tail is described on the hardest IQ test page.

9 What a CRT Score Actually Predicts

The predictive record is genuinely impressive relative to the instrument's length and genuinely small in absolute terms, and both halves belong in the summary. Frederick's Table 5 correlated each cognitive measure with composite indices of decision making behavior built from the time preference and risky choice items in his surveys. The CRT correlated +0.12 with preferring the patient option across 3,099 respondents, +0.22 with choosing a gamble when expected value favored it across 3,150, +0.08 with gambling when expected value favored the sure gain across 1,014, and negative 0.12 with gambling to avoid a sure loss across 1,366.

Those are small correlations. What made them notable is the comparison in the same table. The CRT was either the best or the second best predictor across all four decision making domains and the only measure related to all of them, outperforming instruments running up to 215 items and taking up to three and a half hours. A three item test that predicts as well as a three hour test is a finding about the three hour tests as much as about the three items.

The strongest claim in the literature comes from Toplak, West and Stanovich in Memory and Cognition in 2011, volume 39, pages 1275 to 1289. They administered a battery of 15 classic heuristics and biases problems, covering the conjunction fallacy, base rate neglect, the gambler's fallacy and related tasks, alongside measures of cognitive ability, thinking dispositions and executive functioning. The CRT predicted performance on that battery after all three of those classes of measure had been statistically controlled. Their conclusion was that the CRT captures properties relevant to rational thinking that go beyond what intelligence tests measure.

Read that finding precisely, because it is often stretched. It does not say the CRT is a better measure of intelligence. It says the CRT is not only a measure of intelligence, that some of what it captures is unique, and that the unique part predicts a specific class of reasoning errors. The unique part is the disposition, and a disposition is not an ability. The distinction between a capacity and a tendency to use it, and why the two need different instruments, also runs through the emotional intelligence comparison.

10 What Else Moves a Three Item Score

Because there are only three items, anything that changes one answer changes the score by a third of its range, and several things reliably do. This is the practical cost of extreme brevity, and it shows up clearly in both the original paper and the later meta analytic work.

A sex difference appeared on the CRT that did not appear on any of Frederick's other measures. Men averaged 1.47 on the CRT and women 1.03, with p below 0.0001. On the Wonderlic, administered in the same 45 minute survey under identical conditions, men averaged 26.2 and women 26.5, a difference that was not significant. SAT totals were 1334 and 1324, also not significant. Only SAT math showed a difference, 688 against 666 at p below 0.01, roughly matching the national gap at the time. Frederick also observed that the errors differed in kind: women who missed the widgets item nearly always gave the intuitive answer of 100, while a modest fraction of the men produced other wrong answers such as 20, 500 or 1.

Pablo Branas-Garza, Praveen Kujal and Balint Lenkei tested how far these patterns generalize in a meta study of 118 CRT studies covering 44,558 participants across 21 countries. Their findings were as follows.

  • Being female was negatively associated with the overall score and with each item individually, and the association persisted after controlling for incentives, computerization, student status, and where in the experiment the test appeared.
  • Taking the CRT at the end of an experiment rather than the beginning lowered performance. In their sample, 44.58 percent of studies administered it at the end.
  • Monetary incentives did not improve performance, which is unusual and suggests the failure is not a matter of effort.
  • Student samples outperformed non student samples, and non students were more likely to score zero.
  • Scores drifted upward over the years covered, but the effect was driven by online studies, which is where item exposure concentrates.

The incentive result deserves a second look. Paying people more to get the answer right does not get them the answer, because they do not experience themselves as guessing. They experience themselves as having answered. That is a different failure from insufficient effort and it is the reason the instrument is interesting. Whether education itself moves ability measures, as opposed to moving test familiarity, is examined on the average IQ by education page.

11 How to Read Your Own Three Answers

A CRT result supports one claim about you and no others, and the claim is narrower than the number looks. If you have just answered the three items above for the first time, here is what each outcome does and does not license, stated conservatively.

Your scoreWhat it supportsWhat it does not support
0 of 3You answered at least one item from the first thing that came to mind. In Frederick's pooled sample, 33 percent of respondents did the same.Any statement about your reasoning capacity. The arithmetic in all three items is trivial, so failing them is not evidence about computation.
1 or 2 of 3Mixed performance, which is where the majority sits. Frederick found 28 percent scored one and 23 percent scored two.A meaningful rank between yourself and someone who scored one point differently. One item is a third of the scale.
3 of 3You checked all three. Seventeen percent of Frederick's pooled sample did, rising to 48 percent at MIT.A high ability score. The correlation with tested ability is around .43 to .46, which leaves most of the variance unexplained.
Any score after seeing the items beforeNothing about reflection.Everything. Prior exposure raised scores by d equal to 0.57 in Bialek and Pennycook's pooled data and 0.81 in Haigh's sample.

There is one more caveat that applies to every self administered short measure, including this one. People are not accurate judges of their own cognitive standing, and the CRT is unusual in that failing it feels exactly like passing it. The four lines of evidence in section two make this concrete: respondents who answered 10 cents estimated that 92 percent of others would get the item right, which is the estimate of someone who never noticed there was anything to get wrong. Confidence is not a signal here.

If what you actually want is a number you can place against a population, that requires an instrument with a reference frame, several indicators per construct, and a reported standard error. The page on interpreting a result covers what such a number does and does not mean once you have one, and the score chart shows the bands it falls into.

12 What to Take If You Want an Ability Score

This is not a test of your intelligence, and the short test that is one takes about 45 minutes. That is the whole routing decision on this page, stated plainly, and it follows from everything above rather than from a preference. Three items with no norms, no age adjustment, and known contamination cannot produce an ability estimate. A short battery with multiple indicators per domain can produce a limited one, and it needs roughly three quarters of an hour rather than three minutes.

ACIS is a self-administered online adult assessment covering 20 subtests across the six CHC cognitive domains for ages 16 to 90, sold as a one time payment in three tiers. The Quick form costs 15 dollars, runs six subtests, and takes about 45 minutes. Those six are Similarities and Vocabulary for verbal comprehension, Matrix Reasoning and Figure Weights for fluid reasoning, and Digit Span and Alphanumeric Sequencing for working memory, which returns Verbal Comprehension, Fluid Reasoning and a partial Working Memory index. What the form covers subtest by subtest is on the Quick form page, and the longer options are compared on the adult testing page.

The ACIS reference frame for the published technical figures is documented in the technical manual. The Verbal Comprehension Index reports composite omega of .9745 with a standard error of measurement of 2.40 IQ points and a g loading of .864. The Fluid Reasoning Index reports omega .9727, a standard error of 2.48, and a g loading of .922. Working Memory reports omega .9247, a standard error of 4.12, and a g loading of .788. Those are index level figures for multi subtest composites, which is the level at which reliability of that order is achievable at all. The derivations are in the technical manual.

Two limits belong here rather than in a footer. ACIS is self-administered and is not a clinical instrument, so no result from it should be used for diagnosis, hiring, accommodation requests, or admission to anything. And a self-administered session, however carefully scored, does not substitute for a proctored administration by a licensed psychologist when an institution needs to accept the result. The differences between the two routes are set out on the professional versus online comparison.

13 Sources Behind This Page

Every number on this page comes from one of the following, and each was read rather than cited from a summary. Where a figure appears in a table, the table is named so that it can be checked directly.

  • Frederick, S. (2005). Cognitive Reflection and Decision Making. Journal of Economic Perspectives, 19(4), 25 to 42. Source of the three items in Figure 1, the sample means in Table 1, the correlations with SAT, ACT, Wonderlic and Need For Cognition in Table 4, the decision making correlations in Table 5, and the sex differences in Table 6.
  • Haigh, M. (2016). Has the Standard Cognitive Reflection Test Become a Victim of Its Own Success? Advances in Cognitive Psychology, 12(3), 145 to 149. Source of the 51.4 percent exposure rate and the 2.36 against 1.48 comparison.
  • Stieger, S., and Reips, U. D. (2016). A limitation of the Cognitive Reflection Test: familiarity. PeerJ, 4, e2395. Source of the 44.9 percent exposure rate in a sample of 2,137.
  • Thomson, K. S., and Oppenheimer, D. M. (2016). Investigating an alternate form of the cognitive reflection test. Judgment and Decision Making, 11(1), 99 to 113. Source of CRT-2, the .511 correlation with the original, the disattenuated .905, and the reliability figures.
  • Toplak, M. E., West, R. F., and Stanovich, K. E. (2011). The Cognitive Reflection Test as a predictor of performance on heuristics and biases tasks. Memory and Cognition, 39(7), 1275 to 1289. Source of the claim that the CRT predicts rational thinking performance beyond cognitive ability, thinking dispositions and executive functioning.
  • Toplak, M. E., West, R. F., and Stanovich, K. E. (2014). Assessing miserly information processing: An expansion of the Cognitive Reflection Test. Thinking and Reasoning, 20(2), 147 to 168, doi 10.1080/13546783.2013.844729. Source of the four added items and the .72 reliability of the seven item composite.
  • Primi, C., Morsanyi, K., Chiesi, F., Donati, M. A., and Hamilton, J. (2016). The Development and Testing of a New Version of the Cognitive Reflection Test Applying Item Response Theory. Journal of Behavioral Decision Making, 29(5), 453 to 469. Source of the CRT-Long, the floor effect comparison, the published alpha values from earlier studies, and the correlations with Raven's matrices and numeracy.
  • Bialek, M., and Pennycook, G. (2018). The cognitive reflection test is robust to multiple exposures. Behavior Research Methods, 50, 1953 to 1959. Source of the six study analysis, the d of 0.57, and the 23 Fisher z comparisons.
  • Woike, J. K. (2019). Upon Repeated Reflection: Consequences of Frequent Exposure to the Cognitive Reflection Test for Mechanical Turk Participants. Frontiers in Psychology, 10, 2646. Source of the 65.6 percent and 93.7 percent figures.
  • Branas-Garza, P., Kujal, P., and Lenkei, B. Cognitive Reflection Test: Whom, how, when. MPRA Paper 68049, later published in the Journal of Behavioral and Experimental Economics. Source of the meta study covering 118 studies, 44,558 participants and 21 countries.
  • The Decision Making Individual Differences Inventory entry for the CRT, maintained by the Society for Judgment and Decision Making, which lists the items and the downstream literature.

Two things on this page are not sourced to a study and should be read as such. The step by step solutions in section one are arithmetic, not findings. The interpretation table in section eleven is a reading rule written for this page, built from the published statistics above rather than derived from any single one of them.

The professional framework that governs all of this is explicit. The Standards for Educational and Psychological Testing, published jointly in 2014 by the American Educational Research Association, the American Psychological Association and the National Council on Measurement in Education, require that a score interpretation be supported by evidence for the specific use proposed, that reliability and standard error be reported alongside any score used for decisions, and that the limits of a norm sample be disclosed. The APA standards on test use add the corresponding duty for whoever reports the result: state what the number supports, state what it does not, and do not let a short instrument carry a claim its evidence cannot hold. Applied to the CRT, that is the whole argument of this page in two sentences.

14 Frequently Asked Questions

What is the Cognitive Reflection Test?

A three item measure published by Shane Frederick in the Journal of Economic Perspectives in 2005. Each item has a wrong answer that arrives instantly and a correct answer that requires checking, and Frederick defined what it measures as the ability or disposition to resist reporting the response that first comes to mind.

What are the three questions?

The bat and ball problem, where a bat and ball cost $1.10 and the bat costs $1.00 more than the ball; the machines problem, where 5 machines take 5 minutes to make 5 widgets; and the lily pad problem, where a patch doubles daily and covers a lake in 48 days. All three are printed in Figure 1 of the 2005 paper.

What are the correct answers?

Five cents, five minutes, and 47 days. The answers that arrive first are 10 cents, 100 minutes, and 24 days, and Frederick found those three wrong answers dominated all other wrong answers people gave.

Why is the ball five cents and not ten?

At 10 cents the bat would have to cost $1.10 to be a dollar more, giving a total of $1.20 rather than the stated $1.10. Set the ball at x and the bat at x plus one dollar, and 2x plus 1.00 equals 1.10 gives x equal to 0.05, so the ball is 5 cents and the bat is $1.05.

Why do 100 machines still take five minutes?

Because the fixed quantity is the rate of one machine, not the total. Five machines making five widgets in five minutes means each machine takes five minutes for one widget, so a hundred machines working at the same time each finish one widget in the same five minutes.

Why is the lily pad answer 47 rather than 24?

Because doubling means the day before full coverage is always the day of half coverage, so 48 minus 1 gives 47. On day 24 the patch still has 24 doublings ahead of it and covers a tiny fraction of the lake rather than half.

Is the CRT the world's shortest IQ test?

No. The phrase is an internet label rather than a claim Frederick made. The CRT has no norm sample, no age adjustment, and no conversion to a standard score, so there is nothing for a result to be high or low relative to.

How strongly does the CRT correlate with IQ?

Moderately. Frederick reported .43 with the Wonderlic Personnel Test across 921 respondents, .44 with self reported SAT totals across 434, and .46 with ACT scores across 667. Primi and colleagues reported .32 against Raven's Advanced Progressive Matrices Set I.

What is the average CRT score?

Frederick reported an overall mean of 1.24 out of three across 3,428 respondents in 35 studies. Sample means ranged from 0.57 at the University of Toledo to 2.18 at MIT, which reflects who was recruited rather than a population distribution.

What percentage of people get all three right?

Seventeen percent of Frederick's pooled sample of 3,428. The figure varied enormously by location, from 5 percent at the University of Toledo to 48 percent among the 61 MIT respondents.

Is the CRT just a maths test?

Not entirely, though the numerical component is real. Frederick found the CRT correlated .46 with SAT math and only .24 with SAT verbal, and Campitelli and Gerrans later used mathematical modeling to show the test is not reducible to numeracy.

How reliable is a three item score?

Too low for individual decisions. Frederick did not report a reliability coefficient, and published values from later studies range from about .57 to .74, with Thomson and Oppenheimer reporting .624 for the original three items.

How many people have already seen the items?

Roughly half of a typical sample, and far more online. Haigh found 51.4 percent in 2016, Stieger and Reips found 44.9 percent in a sample of 2,137, and Woike found 65.6 percent among Mechanical Turk participants in 2019.

How much does prior exposure raise the score?

Substantially. Haigh reported means of 2.36 for exposed participants against 1.48 for naive ones, a Cohen's d of 0.81. Bialek and Pennycook found a pooled d of 0.57 across roughly 2,500 participants.

Does that mean the CRT is useless now?

For an individual who has seen the items, yes. For research, Bialek and Pennycook found that in 20 of 23 comparisons the CRT predicted outcomes just as strongly among exposed participants, and in the three exceptions it predicted better rather than worse.

What is the CRT-2?

A four item alternate form published by Thomson and Oppenheimer in 2016, built around linguistic rather than numerical traps. It correlated .511 with the original CRT, which rose to .905 once the unreliability of both short scales was corrected for.

What is the CRT-Long?

A six item version by Primi and colleagues, combining the three original items with three new ones calibrated by item response theory. Its reported alpha was .76, and it cut the proportion of participants scoring zero from 30 percent to 9 percent.

Does the CRT predict anything useful?

Yes, within limits. Frederick found it was the best or second best predictor of time and risk preferences among all his cognitive measures despite having three items, and Toplak, West and Stanovich found it predicted heuristics and biases performance after ability, thinking dispositions and executive functioning were controlled.

Why do men score higher than women on it?

The gap is well documented and not fully explained. Frederick reported 1.47 against 1.03 while finding no significant sex difference on the Wonderlic given in the same session, and a meta study of 118 studies and 44,558 participants found the association held after controlling for incentives, format, and sample type.

Do incentives improve CRT performance?

No. The meta study by Branas-Garza, Kujal and Lenkei found monetary incentives did not change scores, which fits the design: people who answer wrongly do not experience themselves as guessing, so paying them more does not prompt a second look.

What should I take instead if I want an ability score?

Something with a reference frame, several indicators per construct, and a reported standard error. The ACIS Quick form runs six subtests in about 45 minutes for 15 dollars and reports verbal, fluid reasoning and partial working memory indices, and it is a self-administered instrument rather than a clinical or diagnostic one.

Take the assessment

You get a profile, not a number

ACIS measures six CHC domains across 20 subtests and reports each one with its own normed score and confidence interval, so you can see where you are strong and where you are not.

Free trial, no card required. Full report from $15.