Critical Thinking vs Intelligence: Capacity and Use
Intelligence tests estimate cognitive capacity under structured conditions. Critical thinking asks how people evaluate claims, seek evidence, resist bias and revise conclusions. Ability helps, but knowledge, motivation and thinking dispositions determine whether it is applied.
Cognitive capacity supports complex reasoning, while dispositions and knowledge influence whether that capacity is used well.
0 The Short Answer
Thinking well and scoring well are two different achievements, and psychology has spent thirty years accumulating evidence that they can come apart in the same person. An intelligence battery measures computational capacity: how much cognitive machinery you can bring to a problem once you already know a problem is in front of you and that a careful answer is expected. Critical thinking, in the form researchers actually study, asks a different question. Do you notice that the situation calls for deliberation? Do you override the answer that arrives first? Do you hold your own conclusion to the same standard you apply to a conclusion you dislike?
Those questions have empirical answers, and the answers are uncomfortable. In seven studies published in the Journal of Personality and Social Psychology, Keith Stanovich and Richard West (2008) found that susceptibility to several classic reasoning errors, including the conjunction fallacy, anchoring effects and base rate neglect, showed little to no association with measured cognitive ability. Some other tasks in the same package did track ability. The pattern is not "intelligence is irrelevant to reasoning." It is far more interesting than that: intelligence predicts performance on some normative reasoning tasks and is close to useless on others, and which is which follows a logic worth understanding.
Stanovich gave the phenomenon a name in 1993, in the Journal of Learning Disabilities: dysrationalia, the tendency toward irrational thinking and action despite adequate intelligence. The coinage was deliberately provocative, built by analogy with dyslexia, and its point was structural. We have a well developed vocabulary and a well developed test industry for capacity. We have almost neither for the thing capacity is supposed to serve. This page walks through the research that supports the distinction, task by task, and ends with what it means for anyone reading their own score.
What this page is not
This is not an argument that IQ scores are meaningless, nor a pitch for replacing them with a feel good alternative. Cognitive ability is one of the better measured constructs in psychology and it predicts real things. The claim here is narrower and better supported: a high fluid reasoning score certifies capacity, not judgment, and the two are measured by different instruments for good reasons.
1 The Two Questions a Test Can Ask
Stanovich's framework, laid out most fully in Rationality and the Reflective Mind (2011), divides the thinking mind into three levels rather than the usual two. The autonomous mind runs fast, automatic processes that fire whether you want them to or not. The algorithmic mind is the slow, effortful machinery that manipulates symbols, holds intermediate results and executes procedures. The reflective mind decides when the slow machinery gets switched on at all, and what standards the output has to meet before it counts as an answer.
Intelligence tests are excellent instruments for the middle level. That is not a criticism, it is a design specification. A matrix reasoning item hands you a problem, tells you implicitly that there is a single correct completion, gives you a time limit and waits. Every cue that a careful answer is required has already been supplied by the testing situation itself. What remains to be measured is how efficiently your algorithmic machinery runs, and that is exactly what the score reports.
The reflective level is invisible under those conditions, because the test has already done the reflective work for you. Nobody in ordinary life walks up and says "the following statement contains a base rate you are about to ignore, you have ninety seconds." The signal that thought is needed has to be generated internally, and the research question is whether the same individual differences that govern machinery also govern signal detection. The evidence says: only partly, and less than intuition suggests.
This is why the two constructs are separable in principle before you look at a single dataset. Capacity is a ceiling. Rationality is an average across occasions, most of which never announce themselves. A person can have a very high ceiling and a mediocre average, and the tests we use most often are constitutionally unable to notice.
2 Dysrationalia and Why the Word Was Invented
The 1993 coinage was not a joke at anyone's expense. Stanovich's argument was that the learning disability field had a precise concept for a gap between measured ability and a specific skill, and that reasoning deserved the same treatment. Dyslexia describes reading performance that falls well below what general ability would predict. Dysrationalia describes rational thinking that falls well below what general ability would predict. Both are defined by a discrepancy, and a discrepancy is only visible if you measure both sides of it. There is an irony in the borrowed model, because the reading field has since largely retired it: the IQ minus reading discrepancy formula, which for decades decided who counted as dyslexic, was dropped after it proved harmful in practice.
That second measurement is where the field was, and largely still is, undersupplied. Stanovich made the case at book length in What Intelligence Tests Miss (Yale University Press, 2009), which won the 2010 Grawemeyer Award in Education. The book's central complaint is not that batteries are badly built. It is that a construct with enormous cultural weight has been allowed to stand in for a broader idea it was never designed to cover, so that "smart" in ordinary speech quietly means "rational," while the instrument behind the word measures something narrower.
The constructive answer arrived in 2016, when Stanovich, West and Maggie Toplak published The Rationality Quotient: Toward a Test of Rational Thinking with MIT Press, presenting a full assessment framework built from rational thinking tasks rather than ability items. Their reported result is the part worth carrying away: rationality scores correlate with conventional intelligence scores at levels that are low rather than negligible, and on some component tasks the association approaches full dissociation.
Read that carefully, because it cuts both ways. Low positive correlation means the two are not the same variable and cannot be substituted for each other. It also means they are not opposites, and nobody in this literature claims that high scorers reason worse in general. The honest statement is that knowing someone's ability score leaves most of the variance in their rational thinking unexplained, which is precisely the situation in which a second instrument earns its place.
3 Which Biases Track Ability, and Which Do Not
The Stanovich and West paper from 2008 is the sharpest single piece of evidence in this whole area because it does not treat "bias" as one thing. Across seven studies they administered a long menu of reasoning tasks alongside measures of cognitive ability and looked at the associations task by task. The results split cleanly enough to be memorable.
On one side sat errors that were largely independent of ability. Conjunction violations, in which people rate a specific conjoined scenario as more probable than one of its own components. Anchoring, in which an irrelevant number seen moments earlier drags a subsequent numerical estimate toward it. Base rate neglect in several forms. Framing effects, in which logically identical options are chosen differently depending on whether the wording emphasizes gains or losses. High scorers made these errors at rates that were not meaningfully lower than everyone else's.
On the other side sat tasks where ability did predict performance. Belief bias in syllogistic reasoning, where you must judge whether a conclusion follows from premises regardless of whether the conclusion happens to be true in the world, was associated with cognitive ability. So was the choice between probability matching and maximizing in repeated prediction tasks, where the rational strategy is to always pick the more likely outcome rather than distribute guesses in proportion to the base rates.
The organizing principle Stanovich and West proposed is about cues. When the task itself signals loudly that an intuitive response must be suppressed, extra computational capacity gets recruited and helps. When the pull of the intuitive answer is quiet, or when the person never registers that anything needs overriding, capacity sits unused. This is why the split is not random noise across studies. Ability helps when you already know you are in a fight. It does very little when you do not know a fight is happening.
4 The Bias Blind Spot Gets Worse, Not Better
Emily Pronin, Daniel Lin and Lee Ross named the bias blind spot in 2002: the tendency to see cognitive biases operating in other people while remaining convinced you are personally the exception. In a sample of more than 600 United States residents, over 85 percent rated themselves as less biased than the average American, an arithmetic impossibility at the population level and a very ordinary result at the individual one.
The obvious hypothesis is that the blind spot is an information problem. People who know more about cognitive biases, and who have more cognitive resources to apply to self examination, should see themselves more accurately. Richard West, Russell Meserve and Keith Stanovich tested that hypothesis directly in the Journal of Personality and Social Psychology in 2012, under a title that gives away the ending: cognitive sophistication does not attenuate the bias blind spot. Across two studies covering a set of classic biases, none of the blind spots shrank as a function of cognitive ability or thinking dispositions. In several cases the relationship ran the other way, with higher ability accompanying a larger gap between how biased people judged others to be and how biased they judged themselves to be.
Two mechanisms explain the reversal without any mystery. First, self assessment relies on introspection, and biases do not announce themselves in consciousness, so scanning your own mind for evidence of bias returns nothing regardless of how much capacity you devote to the scan. You then read that silence as innocence. Second, greater verbal and reasoning fluency supplies better raw material for justification. A person who can generate three plausible reasons for their existing position in the time it takes someone else to generate one is better equipped to arrive at the conclusion that their position is reasoned rather than motivated.
A separate line of work by Scopelliti and colleagues in 2015 reported that the size of a person's bias blind spot is not related to actual decision making ability, which fits the same picture from another angle. The blind spot is a self perception variable with its own individual differences, and it does not politely follow the ability score around.
5 Wason's Four Cards and the Content Effect
Peter Wason devised the selection task in 1966 and published it in 1968, and it remains the cleanest demonstration that logical competence is not stored in the head as a content free rule. You see four cards. Each has a letter on one side and a number on the other. The visible faces read D, K, 3 and 7. The rule to be tested: if a card has a D on one side, it has a 3 on the other. Which cards must you turn over to find out whether the rule is violated?
The correct selection is D and 7, because only a D without a 3, or a 7 with a D behind it, can falsify the rule. Turning the 3 tells you nothing, since the rule says nothing about what must be on the back of a 3. In Wason's original work fewer than one in ten participants selected the correct pair, and the difficulty replicated for decades afterward. That failure rate is being produced by university students who can state modus tollens correctly when it is written out for them.
Now change the content and leave the logic untouched. Johnson-Laird, Legrenzi and Legrenzi reported in the British Journal of Psychology in 1972 that when the same conditional structure was dressed as a familiar postal rule about envelopes and stamps, 22 of 24 participants produced at least one correct answer with the realistic material, while only 7 of the same 24 managed it with symbolic material. Cox and Griggs followed in Memory and Cognition in 1982 with evidence that direct experience with the specific rule drives the effect, using scenarios like the drinking age regulation that participants had personally policed or been policed by. Cosmides and Tooby argued in 1992 that framing the rule as a social contract, where somebody might be taking a benefit without paying the cost, produces the same jump.
Interpretations of the mechanism still differ across those camps. The finding they share is not in dispute and is what matters here: performance on a logic problem swings enormously with surface content while the logic stays fixed. A fluid reasoning score tells you about the machinery. It does not tell you whether the machinery will be pointed at a problem when the problem arrives wearing everyday clothes.
6 The Cognitive Reflection Test
Shane Frederick published the Cognitive Reflection Test in the Journal of Economic Perspectives in 2005, in an article titled "Cognitive Reflection and Decision Making." It has three items, and they are famous because they are so small.
A bat and a ball cost $1.10 in total. The bat costs $1.00 more than the ball. How much does the ball cost? The answer that arrives instantly is 10 cents. The correct answer is 5 cents.
If 5 machines take 5 minutes to make 5 widgets, how long would 100 machines take to make 100 widgets? The intuitive answer is 100 minutes. The correct answer is 5 minutes.
A patch of lily pads doubles in size every day and covers the lake in 48 days. How long to cover half the lake? The intuitive answer is 24 days. The correct answer is 47 days.
Every one of these requires arithmetic that a competent eleven year old can perform. That is the entire design. Failure cannot be attributed to insufficient computational power, because the computation is trivial once undertaken. What the items detect is whether an answer that feels finished gets audited before it is spoken. Frederick found that CRT performance predicted patterns in time preference and risk preference, and reported differences between groups that a pure arithmetic test would not produce.
The construct held up under harder scrutiny. Toplak, West and Stanovich reported in Memory and Cognition in 2011 that the CRT predicts performance across a broad range of heuristics and biases tasks. In Thinking and Reasoning in 2014 the same team published four additional items to expand the measure, reporting a correlation of .58 between the new set and the original three, an alpha of .72 for the combined seven item version, and continued prediction of rational thinking performance after cognitive ability was accounted for.
That last clause is the crux. If the CRT were simply a short intelligence test, its predictive power would vanish once ability was in the model. It does not vanish. Three arithmetic questions capture something a full ability battery leaves on the table, which is the strongest compact evidence available that reflection is a distinct target of measurement.
7 The Findings in Numbers
Arguments in this area drift toward the abstract very quickly, so here are the six results that carry the most weight, each attached to a published source. These are the figures worth remembering when someone insists that a high score settles the question of whether a person reasons well.
Under 10%
Share of participants selecting the correct two cards in Wason's original abstract selection task, a result replicated for decades with educated samples.
22 of 24 vs 7 of 24
Participants producing at least one correct answer with realistic versus symbolic material in Johnson-Laird, Legrenzi and Legrenzi (1972). Identical logic, different clothing.
7 studies
The size of the Stanovich and West (2008) package showing conjunction, anchoring and base rate errors largely independent of cognitive ability.
85%+
Share of over 600 US residents in Pronin, Lin and Ross (2002) who rated themselves as less biased than the average American.
alpha .72
Reliability of the seven item cognitive reflection measure in Toplak, West and Stanovich (2014), which still predicted rational thinking with ability controlled.
3 items
The length of Frederick's (2005) original test, every item solvable with grade school arithmetic, yet routinely failed by university students.
Notice that none of these findings requires a fringe theory to interpret. They come from mainstream journals, they have been replicated in various forms, and they were mostly produced by researchers who take cognitive ability seriously as a construct. The dissociation between capacity and judgment is not a heterodox position. It is the ordinary reading of the data by the people who collected it.
8 When Ability Argues Back
The most unsettling strand of this literature concerns what happens when reasoning has a stake in the outcome. Dan Kahan's 2013 paper in Judgment and Decision Making examined the relationship between cognitive reflection and ideologically charged conclusions, and found that greater reflection did not reliably move people toward the position better supported by evidence. It moved them toward positions congenial to their group, more consistently. That is one reason the literature gathered in IQ and political orientation stays so hedged: the associations it reports are small, they differ depending on whether social or economic attitudes are measured, and the group distributions overlap far more than they diverge.
The 2017 study by Kahan, Peters, Dawson and Slovic in Behavioural Public Policy made the mechanism concrete with a single elegant manipulation. Participants were given a two by two table of outcome data and asked which conclusion it supported. For half of them the table described a skin cream trial. For the other half the identical numbers described a city's handgun ban. On the skin cream version, accuracy rose with numeracy exactly as you would expect: better quantitative skill, better answer. On the gun version, accuracy depended on whether the correct answer flattered the participant's politics, and the polarization was largest among the most numerate participants.
Read the shape of that result carefully. Quantitative ability did not fail. It performed beautifully, in service of a conclusion selected before the analysis began. The skill functioned as an amplifier rather than a corrective, which means that under motivated conditions the relationship between capacity and accuracy can invert. This is the same pattern the bias blind spot work found from a different direction, and it recurs often enough that it deserves a place in anyone's mental model of what a high score does and does not buy.
The practical lesson is unglamorous. On neutral material, capacity helps. On material where you already want a particular answer, capacity is not neutral equipment, and the protection has to come from somewhere else: a procedure, an adversary, a pre commitment, a number written down before the data arrives. Nothing in the ability score supplies any of those.
9 What Critical Thinking Tests Actually Measure
Critical thinking assessments exist as a separate industry, and comparing their contents with an ability battery makes the distinction concrete. The Watson-Glaser Critical Thinking Appraisal, the oldest instrument still in wide commercial use, is organized around five operations: drawing inferences from stated facts, recognizing unstated assumptions, evaluating deductive validity, interpreting whether conclusions follow from evidence, and judging the strength of arguments. Items are built from ordinary prose passages, frequently on contested topics, and the examinee's job is to separate what the passage supports from what it merely suggests. Watson-Glaser is no laboratory curiosity either; it sits inside the standard employer testing lineup next to short general ability screeners like the Wonderlic, and it turns up most often in law and management hiring.
The Halpern Critical Thinking Assessment, developed by Diane Halpern, takes a different route by pairing open ended responses with forced choice items on the same scenarios, so that recognition of the right answer and spontaneous production of it are scored separately. Heather Butler reported in Applied Cognitive Psychology in 2012 that scores on this instrument predicted real world outcomes across domains including education, health, legal matters, finance and relationships, measured as the frequency of negative life events attributable to poor judgment.
Butler, Pentoney and Bong followed in Thinking Skills and Creativity in 2017 with a direct comparison, concluding that critical thinking ability was the better predictor of consequential life decisions than intelligence in their sample. Treat that as one study rather than a settled law, and note the boundary conditions: outcome measures of this kind depend on self report, and critical thinking scores and ability scores correlate substantially, so the instruments are cousins, not strangers.
Now set that against a cognitive battery. A battery samples verbal comprehension, fluid reasoning, visual spatial processing, working memory, processing speed and quantitative reasoning, and reports an index for each. Not one of those indices asks whether you accept a conclusion on insufficient evidence, because accepting conclusions is not what the items are about. The two families of tests are aimed at different targets, and if you want to know something about a person's judgment, the family with judgment in its item content is the one to consult. If you want to understand what the ability side covers in detail, our page on what IQ tests actually measure lays out each domain.
10 Dispositions Are Not Capacity
The third component in this picture is the one people find hardest to accept as measurable: what a person is inclined to do, as opposed to what they are able to do. The research calls these thinking dispositions, and they are assessed by questionnaire rather than by performance items.
Actively open-minded thinking. The willingness to seek evidence against a favored belief and to weight it fairly when it appears. Stanovich and West reported in the Journal of Educational Psychology in 1997, with 349 students, that this disposition predicted the ability to evaluate arguments independently of prior belief even after cognitive ability had been partialled out of the analysis.
Need for cognition. The tendency to seek out and enjoy effortful thought, introduced by John Cacioppo and Richard Petty in the Journal of Personality and Social Psychology in 1982 and now one of the most heavily used scales in the field. Two people with identical reasoning capacity can differ enormously in how often they choose to deploy it.
Tolerance of ambiguity. The capacity to leave a question open rather than reaching for premature closure. Impatience for resolution produces confident errors, and it is not correlated with lacking the ability to work the problem out.
Dispositions map onto the reflective level of Stanovich's architecture rather than the algorithmic level, which is why they add predictive power after ability is controlled. They also carry a real methodological caveat that honest coverage should state plainly: self report scales measure self perception, and people who consider themselves open minded are precisely the population you would expect to have trouble noticing when they are not. That limitation is one reason performance based rational thinking tasks, of the sort assembled in the 2016 Stanovich, West and Toplak framework, matter as a complement.
The compact way to hold all three levels together: capacity sets what you can do at your best, dispositions determine how often your best is actually attempted, and knowledge of the relevant rules determines whether the attempt lands. A test of any one of the three tells you little about the other two.
11 Reading Your Own Score Honestly
Suppose you take a full battery and the fluid reasoning index comes back high. What have you learned, precisely? You have learned that on a set of novel problems, under time pressure, with no relevant background knowledge required and no motivational stake in the answer, your inference machinery ran efficiently relative to the norm sample. That is a genuine finding about a genuine capacity, and it is worth knowing. What it is not is a label for a kind of mind, and that is the quiet error underneath the popular list of intelligence types: Gardner proposed in 1983 that the intelligences were largely independent, and the abilities researchers actually measure turn out to correlate instead.
What you have not learned is whether you audit conclusions you like, whether you notice base rates when nobody points at them, whether you can be argued out of a position by evidence, or whether your quantitative skill would serve accuracy on a question where your identity is involved. Those are separate measurements, and the score is silent on all of them. It is not evidence against them either, which is worth saying clearly. It is silence.
The same reading discipline applies across the profile. A strong verbal comprehension index reflects accumulated vocabulary and verbal knowledge, which is a different developmental story from novel problem solving, a distinction covered in more depth in our comparison of fluid and crystallized intelligence. A strong working memory index says something about holding and manipulating information, not about what you choose to hold it for. Interpreting a profile means reading each index for what its items sample, then stopping.
ACIS is an online self assessment rather than a clinical evaluation, and it reports a profile across six CHC domains with confidence intervals around each index rather than a single naked number. That structure exists because a composite hides exactly the information that makes a profile interesting: which domains are relatively stronger, and how wide the uncertainty is around each estimate. If you want to see where any particular index falls relative to the population, the IQ percentile chart converts standard scores into rarity, and our discussion of what counts as a good IQ score covers how those thresholds are conventionally read.
12 How Good Judgment Is Actually Trained
Capacity is substantially constrained by the time you are an adult. The reflective layer is not, and everything discussed above suggests where the leverage sits. The interventions with the best support share a family resemblance: they replace an in the moment virtue with a procedure that runs whether or not you feel like being fair.
Charles Lord, Mark Lepper and Elizabeth Preston demonstrated the archetype in the Journal of Personality and Social Psychology in 1984. Asking participants to consider the opposite, to actively construct the case for the conclusion they did not reach, reduced biased judgment where generic instructions to be fair and unbiased did not. The specificity is the whole point. A general intention to think well changes very little. A concrete operation with a defined output changes something.
Write the falsifier first. Before you evaluate evidence, state what result would move you off your current position. If nothing would, you are not reasoning about the question, and you have learned that in advance rather than after an argument.
Learn the specific rules, not the general attitude. Base rates, regression to the mean, sample size, conditional probability and the difference between confirming and falsifying a conditional are teachable content. The Wason results show that recognizing the structure in unfamiliar packaging is the hard part, and packaging variety is what practice supplies.
Slow down where the answer felt free. The bat and ball item is a portable diagnostic for a general habit. When an answer arrives with no sensation of effort, that absence of effort is the cue to audit, and it is the one cue you can actually detect from the inside.
Get an external check on motivated ground. On questions tied to identity, group or self interest, the motivated numeracy results say your own skill is not a safeguard. Build the check into the process: a named critic, a pre registered prediction, a number committed to paper before the data lands.
None of this raises a fluid reasoning index, and it is not supposed to. Ability scores and reflective habits respond to different inputs, which is the practical restatement of everything on this page. Related reading: how ability scores behave as predictors in employment settings is covered in our page on IQ and job performance, and the neighboring question of whether a separate affective capability belongs in the same conversation is handled in our treatment of emotional intelligence versus IQ.
If you want a measured starting point for the capacity side, that part is straightforward to obtain. Knowing where your reasoning machinery sits is useful precisely because it stops you from confusing it with the other thing.
Are critical thinking and intelligence the same thing?
No. They overlap, but published assessments of rational thinking correlate only modestly with ability batteries, leaving most of the variance in each unexplained by the other.
Can a person with a very high IQ reason badly?
Yes, and the pattern has a name. Stanovich labeled it dysrationalia in 1993 to describe reasoning that falls short of what general ability would predict.
Does a high score protect against cognitive bias?
Against some, mildly. Against conjunction errors, anchoring and base rate neglect, the 2008 evidence found essentially no protective effect from higher ability.
What is the bias blind spot?
The habit of spotting distortions in other people's thinking while judging your own reasoning to be clear. Pronin and colleagues documented it in 2002.
Is it true that smarter people have a bigger blind spot?
In the 2012 West, Meserve and Stanovich data, yes, in several comparisons. Cognitive sophistication never shrank the gap and sometimes widened it.
Why does introspection fail here?
Biased processing leaves no trace in awareness. Searching your own mind returns nothing, and the emptiness gets misread as proof of impartiality.
What is the Cognitive Reflection Test?
A three question measure Frederick published in 2005, built so that each item has a tempting wrong answer and a correct one requiring almost no arithmetic.
Is the CRT just a short IQ test?
No. It shares variance with ability, but it continues to predict rational thinking performance in analyses where ability has already been accounted for.
Why is the bat and ball answer five cents?
If the ball were ten cents the bat would be a dollar ten, giving a total of one twenty. Five and one oh five sum correctly to one ten.
Does failing the CRT mean someone is unintelligent?
No. Plenty of high scoring people miss items, which is the finding that makes the instrument interesting rather than redundant.
What does the Wason selection task show?
That people rarely look for the case that would disprove a conditional rule, even when they can state the underlying logic when asked directly.
Why does familiar content make the task easier?
Researchers disagree on the mechanism, citing pragmatic schemas, direct experience with the rule, or evolved cheater detection. The performance difference itself is robust.
What is motivated reasoning?
Processing evidence in the direction of a conclusion you already prefer, while experiencing the process as ordinary careful analysis.
Can quantitative skill make political bias worse?
The 2017 Kahan and colleagues study found the sharpest partisan divergence among participants with the strongest numerical skills on politically framed data.
What do thinking dispositions add to prediction?
Incremental variance. Questionnaire measures of open minded thinking forecast reasoning quality beyond what an ability score alone accounts for.
What is need for cognition?
A stable preference for effortful thinking, introduced by Cacioppo and Petty in 1982. It concerns appetite for mental work, not the amount available.
Are self report thinking scales trustworthy?
Partially. They predict useful outcomes, but people rate their own open mindedness generously, which is why performance tasks belong alongside them.
Which tests assess critical thinking directly?
The Watson-Glaser appraisal and the Halpern assessment are the two most established, both built from prose scenarios rather than abstract puzzle items.
Can critical thinking be improved through training?
Specific techniques with defined steps show better results than broad exhortations to be objective. Lord and colleagues demonstrated this contrast in 1984.
Does an ACIS report include a rationality measure?
No. It covers six CHC ability domains with intervals around each index, and makes no claim about judgment quality or decision habits.
If judgment matters more, why measure ability at all?
Because the two answer different questions and one does not replace the other. Capacity data is informative on its own terms once you stop asking it to certify wisdom.
14 Sources Behind This Page
The claims on this page follow the published literature rather than our own assertions. These are the primary papers and reference bodies worth reading directly, with what each contributes.
Nisbett, R.E. et al. (2012). Intelligence: new findings and theoretical developments. American Psychologist, 67(2). The broad APA review of what moves measured intelligence and what does not.
Gottfredson, L.S. (1997). Mainstream Science on Intelligence, the editorial signed by 52 researchers. Intelligence, 24(1), hosted by the University of Delaware. A consensus statement on what IQ tests measure.
Spearman, C. (1904). General intelligence, objectively determined and measured. American Journal of Psychology, full text at Classics in the History of Psychology. The paper where the g factor entered psychology.
Plomin, R. & Deary, I.J. (2015). Genetics and intelligence differences: five special findings. Molecular Psychiatry. Open access. The standard modern review of what twin and DNA evidence does and does not show.
Voncken, L., Albers, C.J. & Timmerman, M.E. (2019). Improving confidence intervals for normed test scores. Behavior Research Methods. Open access. Documents the mean 100, SD 15 metric and the uncertainty that norming from samples adds to any score.
Pearson Clinical Assessment Scientific Council (2023). Standardized Clinical Assessment for Practitioners: A Primer. How standard scores, percentile ranks and the standard error of measurement are meant to be read together.
Pearson (2008). WAIS-IV Score Report sample. What a real report contains: every composite paired with a percentile rank, a 95% confidence interval and a qualitative description, never a bare number.
Pearson (2024). WAIS-5, Wechsler Adult Intelligence Scale, Fifth Edition. The current adult battery, covering ages 16:0 to 90:11 across five cognitive domains.
Riverside Insights. Stanford-Binet Intelligence Scales, Fifth Edition. A second current publisher with a different structure and a different age span, from 2 to 85 and above.
Take the assessment
You get a profile, not a number
ACIS measures six CHC domains across 20 subtests and reports each one with its own normed score and confidence interval, so you can see where you are strong and where you are not.