The fluid intelligence test and the four tasks that measure it
A fluid reasoning test is a set of problems whose rules are not stated and cannot be looked up, arranged so that the only route to an answer is to work the rule out from the material in front of you. Four task families do almost all of the work: matrices, series completion, quantitative balance problems, and analogies. This page covers what each one asks, what a score reports, and where the evidence is contested.
Almost every problem in the Raven matrices can be described by five rule types, and the count of rules in a problem predicted more than half the variance in how often it was failed.
0 Quick Answer
A fluid intelligence test measures reasoning on material the test taker has not seen before, using tasks where the rule is discoverable from the item itself and prior knowledge gives almost no advantage. The construct is Gf in the Cattell-Horn-Carroll framework, and in the current taxonomy it decomposes into three narrow abilities. Schneider and McGrew list them in their 2018 chapter in Contemporary Intellectual Assessment, fourth edition, pages 73 to 163, published by Guilford Press: Induction, the ability to observe a phenomenon and discover the underlying principles that determine its behavior; General Sequential Reasoning, the ability to reason logically from known premises; and Quantitative Reasoning, the ability to reason with quantities and mathematical relations.
In practice that translates into four task families. Matrices, where a grid of figures follows rules that have to be induced and applied to a missing cell. Series completion, where a sequence of numbers or figures continues under a rule that is never named. Quantitative balance problems, where unstated equivalences between objects have to be chained together. Analogies, verbal or figural, where a relation demonstrated in one pair has to be transferred to another. Every widely used fluid reasoning measure is built from some combination of these.
Gf matters disproportionately because it carries more of the general factor than any other broad domain. In the ACIS technical manual's normative model, the Fluid Reasoning Index has a g loading of .922, the highest of the six primary indices, above Visual Spatial at .906, Quantitative Reasoning at .882, Verbal Comprehension at .864, Working Memory at .788, and Processing Speed at .648. How Gf differs from accumulated knowledge is the subject of the fluid versus crystallized comparison rather than this page, which is about the tasks and the score.
.922
The g loading of the ACIS Fluid Reasoning Index, the highest of the six primary indices in the published technical manual.
57 percent
Of the variance in mean error rates across 32 Raven problems explained by the number of rules in the problem, in Carpenter, Just, and Shell's 1990 analysis in Psychological Review.
.05
The far transfer effect of working memory training on nonverbal ability against treated controls, across 67 comparisons, in Melby-Lervag, Redick, and Hulme's 2016 meta-analysis.
A task qualifies as a fluid reasoning measure when the rule needed to answer it is present in the item and absent from the test taker's education. That single criterion does most of the design work. It rules out vocabulary, general information, and arithmetic facts, because those can be studied. It rules out anything where a taught procedure produces the answer directly. What remains is a class of problems where the material is arbitrary enough that nobody has an advantage from having met it before, and structured enough that a rule can be recovered from it.
The second design requirement is that the rule be recoverable within the item rather than across the test. A good fluid item is self contained. Everything needed to justify one answer and reject the others is visible, which is what makes it possible to explain afterward why the key is the key. That property also makes fluid items unusually easy to write badly, because a stem that admits two defensible readings produces an item where the strongest reasoners split.
The third requirement is graded difficulty. A fluid reasoning subtest that is uniformly easy measures nothing above the middle, and one that is uniformly hard measures nothing below the top. The items are ordered so that early ones establish that the format is understood and later ones stack demands. In the ACIS battery each subtest discontinues after three consecutive errors or timeouts, so a test taker stops when the items have moved past their range rather than working through a tail of problems that carry no information about them. Why difficulty behaves that way is worked through on the page on test difficulty and ceilings.
What a fluid reasoning test is not is a test of speed. Some fluid tasks are timed for practical reasons, but the construct is power rather than rate, and processing speed is reported separately for exactly that reason. The distinction between a time limit as an administrative convenience and a time limit as part of the construct is set out on the page on time limits.
2 Where the Construct Came From
The distinction between reasoning and knowledge was proposed by Raymond Cattell in the 1940s and only tested properly two decades later. Cattell set out the idea in The Measurement of Adult Intelligence, Psychological Bulletin, 40(3), 153 to 193, published in 1943. Carroll, reviewing the record in 1984, could not find the theory stated anywhere in Cattell's writing before that paper, so 1943 is the earliest defensible date for the construct rather than the 1941 conference abstract that is sometimes cited.
The empirical test came in 1963. Cattell published Theory of Fluid and Crystallized Intelligence: A Critical Experiment in the Journal of Educational Psychology, 54(1), 1 to 22, factoring results from culturally embedded and culture fair tests in a sample of 277 seventh and eighth grade children. Two general ability factors emerged rather than one, the first aligning with traditional intelligence tests and the second with the culture reduced measures. That two factor result is what turned a proposal into a research program. He developed it at book length in Abilities: Their Structure, Growth, and Action, published by Houghton Mifflin in 1971, and revised it under a new title, Intelligence: Its Structure, Growth, and Action, published by North-Holland in 1987.
The framework used by test publishers today came from a third source. John Carroll's Human Cognitive Abilities: A Survey of Factor-Analytic Studies, published by Cambridge University Press in 1993, reanalyzed more than 460 data sets from the factor analytic literature and produced a three stratum model with a general factor at the top, broad abilities beneath it, and narrow abilities beneath those. Merging Carroll's stratum structure with Cattell and Horn's domain list produced the Cattell-Horn-Carroll model that the CHC page describes.
What changed in the most recent revisionThe narrow ability list under Gf has shrunk. McGrew's 2009 review in Intelligence, 37(1), 1 to 10, listed five narrow abilities under Gf, including Piagetian Reasoning and Speed of Reasoning alongside Induction, General Sequential Reasoning, and Quantitative Reasoning. The 2018 Schneider and McGrew chapter lists three. Anyone citing a five ability list under Gf is citing the older taxonomy.
3 The Four Task Families, With Worked Examples
Almost every fluid reasoning item in professional use belongs to one of four families, and each is worth walking through in words. None of the examples below is an ACIS item. They are written here to show the reasoning, not to preview a subtest.
Matrices. Picture a three by three grid. The top row holds a rounded square that is plain, then the same square lightly dotted, then the same square with diagonal stripes. The middle row holds a star, plain, dotted, striped. The bottom row holds a hexagon, plain, dotted, and then an empty cell. Two rules are running at once. Shape is constant across each row and changes down the column. Fill pattern is distributed across each row in the order plain, dotted, striped. Both rules point at the same answer for the missing cell, a striped hexagon, and the wrong options are the ones that satisfy one rule while violating the other.
Series completion. Take the sequence 2, 6, 12, 20, 30. Nothing about it is memorized. Subtract each term from the next and you get 4, 6, 8, 10, a constant second difference of two. That means the next gap is 12 and the next term is 42. The reasoning is two layered: notice that the first differences are not constant, then treat those differences as a series in their own right. A series item gets harder not by using larger numbers but by adding a layer, alternating two interleaved rules, or making the operation change position within the term.
Quantitative balance problems. Suppose two circles balance three triangles, and one triangle balances two squares. How many squares balance four circles? Chain the equivalences. Four circles equal six triangles by doubling the first statement. Six triangles equal 12 squares by the second. The answer is 12. Nothing here requires arithmetic beyond multiplication, and the difficulty lives entirely in holding two unstated equivalences together long enough to compose them.
Analogies. A figural analogy shows a shape, then the same shape rotated a quarter turn with a dot added, then a second shape, and asks for the fourth term. The relation demonstrated in the first pair, rotate and add, is applied to the third term. Verbal analogies use the same structure over words, and they are the one family that sits partly outside Gf, because recognizing the relation between two words requires knowing both. That is why analogy formats appear in verbal comprehension measures as well as fluid ones, and why a pure Gf battery leans on the figural version.
Task family
Narrow ability it leans on
What raises difficulty
Contamination risk
Matrices
Induction
More rules operating at once, and harder correspondence between elements
Low. Perceptual strategies work only on the easiest items
Moderate for numeric series, since arithmetic fluency helps
Quantitative balance
Quantitative Reasoning
Longer chains of equivalence and more objects to track
Low, provided the arithmetic stays trivial
Figural analogies
Induction
Multiple simultaneous transformations, including reflection
Low
Verbal analogies
Induction, but over learned material
Rarer words and more abstract relations
High. Word knowledge is doing part of the work
4 What a Matrix Item Actually Tests
The best evidence about what these tasks measure comes from taking one apart item by item, which Carnegie Mellon researchers did in 1990. Patricia Carpenter, Marcel Just, and Peter Shell published What One Intelligence Test Measures in Psychological Review, 97(3), 404 to 431. They recorded eye fixations and think aloud protocols from 12 Carnegie Mellon students working Raven problems, collected rule descriptions from a further 22 students across two sessions, and built two computer simulations, FAIRAVEN and BETTERAVEN, that reproduced the performance of the median and the best performers respectively.
Their first result was a taxonomy. Across the problems they examined, five rule types governed the variation: constant in a row, quantitative pairwise progression, figure addition or subtraction, distribution of three values, and distribution of two values. Nearly every problem in Sets I and II can be classified with respect to which of those govern it. Verguts and De Boeck reported the same five types independently in the European Journal of Cognitive Psychology, 14(4), 521 to 547, in 2002.
Their second result explains difficulty. A simple linear regression with a single predictor, the number of rules in a problem, "accounted for 57% of the variance among the mean error rates" across the 32 problems they had classified. Their error rates correlated .91 with the rates Forbes had published in 1964 from 2,256 British adults, so the pattern was not a quirk of a student sample.
Their third result is the one that reaches beyond the Raven. The processes distinguishing high from low performers were, in their words, "primarily the ability to induce abstract relations and the ability to dynamically manage a large set of problem-solving goals in working memory." Extra rules do not make each inference harder. They make the plan harder to hold. In a separate experiment with 45 students, Raven errors correlated .77 with errors on the Tower of Hanoi, a task with no figural content at all, which is difficult to explain unless something like goal management is shared between them.
That finding is why the working memory question in a later section is a real dispute rather than a quibble, and why a fluid reasoning battery that also measures working memory separately gives you more information than one that does not. The ACIS working memory subtests are described on the cognitive domains page.
5 The Named Instruments That Measure Gf
Four instruments account for most fluid reasoning measurement, and they differ in what they return more than in what they ask. Knowing which one produced a number tells you what the number can be compared against.
Instrument
Task families used
What it reports
What it cannot tell you
Raven's Advanced Progressive Matrices
Matrices only
A raw score out of 36 in Set II, expressed as a percentile against the reference group used
Anything about verbal ability, working memory, speed, or the shape of a profile
Raven's Progressive Matrices 2, published by Pearson in 2018
Matrices only
A standardized score for ages 4 to 90, delivered digitally or on paper
The same limits, plus the item count and administration time are not published on Pearson's public pages
Wechsler adult and child scales
Matrices, quantitative balance, and figure classification
A Fluid Reasoning Index alongside four other index scores and a Full Scale IQ
Narrow ability separation within Gf, since the index rests on two or three subtests
Culture reduced batteries in the Cattell tradition
Matrices, series, classification, and conditions
A single composite intended to minimize language and schooling effects
A domain profile, which is the point of the design rather than a defect
ACIS Fluid Reasoning Index
Matrices, series, quantitative balance, constraint reasoning, and relational integration
An index from five subtests with an omega of .9727 and a standard error of measurement of 2.48 points
A clinical diagnosis, a hiring recommendation, or an accommodation decision
The Raven family is the canonical measure for a reason. It was built by John Raven, a student of Spearman, around a single format that requires no reading and no instruction beyond a demonstration, which is what made it usable across languages and education levels. Carpenter, Just, and Shell describe the item set as running from problems solvable by simple perceptual continuation to problems that cannot be, and that range within one format is unusual. The current edition and what it covers are on the Raven's 2 page.
Its limits follow from the same design. A matrices only instrument returns one number and no profile, so it cannot say whether a low result reflects weak induction, a working memory constraint, or a strategy that stalled. It also concentrates the entire measurement on one item format, which means a person unusually good or unusually bad at that specific format has no counterweight. Instruments that reduce language deliberately are compared on the culture fair page and the wider format question is on the nonverbal test page.
6 Why Fluid Reasoning Carries So Much of the General Factor
Fluid reasoning sits closer to the general factor than any other broad domain, and the strongest version of that claim is contested by the data that built the framework. Both halves of that sentence matter, so take them in order.
The strong position comes from Jan-Eric Gustafsson. In A Unifying Model for the Structure of Intellectual Abilities, Intelligence, 8(3), 179 to 203, published in 1984, he fitted a hierarchical model to a battery of tests and concluded that the second order factor of fluid intelligence is identical with the third order general factor. If that is right, Gf is not merely the most g loaded domain. It is g, measured through a particular set of tasks.
Later work put a number on it under specified conditions. Valentin Kvist and Gustafsson published The Relation Between Fluid Intelligence and the General Factor as a Function of Cultural Background in Intelligence, 36(5), 422 to 436, in 2008. Fitting a second order model to 17 tests across three groups, 2,358 Swedes, 620 European immigrants, and 591 non European immigrants, they reported a g to Gf relationship of .83 in the combined group and close to unity within each of the three subgroups considered separately. Their reading is that pooling groups with different learning opportunities is what pulls the relationship below one.
The counterweight comes from the Carroll data itself. Kevin McGrew, reviewing the three stratum model thirty years on in the Journal of Intelligence, 11(2), 32, in 2023, reports that in Carroll's own analyses the Gf loading on g was .83, similar to the loading for comprehension knowledge at .83, and did not approach unity. Treating Gf and g as the same thing is therefore a defended position rather than a settled one, and this page states it as such.
The ACIS figures are consistent with the weaker and better supported claim. The Fluid Reasoning Index loads .922 on g, the highest of the six primary indices, with Visual Spatial at .906 and Verbal Comprehension at .864. Both the General Ability Index at .954 and the reduced verbal composite at .949 load higher than any single domain, which is what you expect if g is best estimated from breadth rather than from any one family of tasks. The statistical core of the general factor is on the g factor page and what a composite is built from is on the page on what IQ measures.
7 The Five Subtests Behind the ACIS Fluid Reasoning Index
An index built from five indicators behaves differently from one built from two, and the ACIS Fluid Reasoning Index uses five deliberately. Each covers a different route into the same construct, which is what stops the index from becoming a measure of one item format.
Logic Grid, g loading .884. A constraint satisfaction problem where several conditions have to be combined to produce the only arrangement that satisfies all of them. This is General Sequential Reasoning in the Schneider and McGrew sense, deduction from stated premises, and it is the highest loading fluid subtest in the battery.
Figure Weights, g loading .872. Quantitative balance problems of the kind worked through earlier. Unstated equivalences between objects have to be chained, which places it under Quantitative Reasoning while keeping the arithmetic itself trivial.
Complex Relations, g loading .856. Several relations have to be integrated at once rather than applied in sequence. This is the subtest most sensitive to the goal management demand that Carpenter, Just, and Shell identified.
Matrix Reasoning, g loading .840. The classic induction format, a grid governed by rules that have to be recovered and applied to a missing cell. The most widely recognized fluid task and the one with the longest research record behind it.
Visual Number Series, g loading .837. Sequence completion where the generating rule is never named. Induction over numeric material, with the layered difficulty structure described in the worked example above.
Notice that the ordering does not follow how hard the tasks feel. Logic Grid carries more of the general factor than Complex Relations even though most test takers describe the latter as harder to explain afterward. Subjective difficulty and psychometric information are separate axes, which is the whole argument of the page on what difficulty buys.
.9727
Omega reliability of the ACIS Fluid Reasoning Index across its five subtests, per the published technical manual.
2.48 points
The standard error of measurement on the Fluid Reasoning Index, against 5.22 points on the two subtest Processing Speed Index.
.813
The correlation between the Fluid Reasoning Index and the Visual Spatial Index, the highest of the FRI intercorrelations, with Quantitative Reasoning at .812 and Verbal Comprehension at .801.
The reliability difference between those two indices is worth reading carefully, because it is mostly a count. Five indicators average out more measurement noise than two do, which is why the standard error of measurement more than doubles when a domain is carried by a shorter subtest set. What each subtest format is for across the whole battery is set out on the subtest formats page.
8 The Working Memory Dispute, Presented as a Dispute
How much of fluid reasoning is working memory is an open question with two published positions, both from serious groups, and the numbers are far apart. The exchange ran in a single issue of Psychological Bulletin in January 2005 and has not been closed since.
Phillip Ackerman, Margaret Beier, and Mary Boyle argued for separation in Working Memory and Intelligence: The Same or Different Constructs?, Psychological Bulletin, 131(1), 30 to 60. They ran a meta-analysis of 86 samples relating working memory to intelligence and reported that "the average correlation between true-score estimates of WM and g is substantially less than unity," at .479. Their conclusion was that the two constructs overlap substantially and are not the same thing.
Michael Kane, David Hambrick, and Andrew Conway replied in the same issue in Working Memory Capacity and Fluid Intelligence Are Strongly Related Constructs, Psychological Bulletin, 131(1), 66 to 71. They agreed that working memory capacity is not isomorphic with Gf, but argued that Ackerman and colleagues had understated the association by working with correlations between individual tasks rather than between latent variables. Their reanalysis of 14 data sets from 10 published studies, representing more than 3,100 young adult subjects, produced a median correlation of .72 between working memory capacity and Gf, which they read as roughly 50 percent shared variance.
A third comment ran alongside it. Klaus Oberauer, Ralf Schulze, Oliver Wilhelm, and Heinz-Martin Suss published a separate critique at pages 61 to 65 of the same issue, arguing that reanalysis with correct statistical procedures shows g and working memory capacity to be very highly correlated. Beier and Ackerman replied at pages 72 to 75 under the title Working Memory and Intelligence: Different Constructs.
Why the gap between .479 and .72 is not a contradictionThe two numbers answer different questions. A correlation between observed task scores includes each task's measurement error and its task specific variance, both of which pull the estimate down. A correlation between latent factors, each defined by several tasks, removes that error and estimates the relationship between the constructs. Neither figure is wrong. The dispute is about which one describes what a practitioner should expect, and that has never been settled by a decisive study.
The practical consequence for reading a score is straightforward. A fluid reasoning index and a working memory index will usually move together, and in the ACIS normative model the correlations across the six primary indices reflect that structure, with Working Memory to Visual Spatial at .715 and Working Memory to Processing Speed at .475. If someone's Gf is far below their Gwm, or the reverse, that gap is information about how the two are working in that person rather than evidence that one of the tests failed. What working memory contributes to reasoning in ordinary tasks is discussed on the IQ and memory page.
9 Whether Fluid Intelligence Can Be Trained
The claim that a working memory exercise raises fluid intelligence was made in 2008, tested properly in 2013, and has not survived meta-analysis. This is one of the cleanest examples in psychology of a finding that replicated in the press and not in the design.
Susanne Jaeggi, Martin Buschkuehl, John Jonides, and Walter Perrig published Improving Fluid Intelligence With Training on Working Memory in the Proceedings of the National Academy of Sciences, 105(19), 6829 to 6833, in 2008. Seventy young adults took part, 34 trained on an adaptive dual n-back task for 8, 12, 17, or 19 days and 35 served as controls, with Raven's Advanced Progressive Matrices and the Bochumer Matrizen-Test as outcome measures. Trained participants improved more than controls, and the authors reported that gains scaled with training time, describing the effect as dosage dependent.
The design had two weaknesses that later work targeted. The control group was not given an alternative task, so expectancy and engagement were not held constant, and the training cells were very small. Redick and colleagues note that the four groups contained roughly seven, eight, or eleven participants each.
Thomas Redick and seven co-authors ran the corrected version. No Evidence of Intelligence Improvement After Working Memory Training appeared in the Journal of Experimental Psychology: General, 142(2), 359 to 379, in 2013. Seventy-five participants completed all sessions and 73 entered the analysis, split into 24 doing 20 sessions of adaptive dual n-back, 29 doing an adaptive visual search task as an active placebo, and 20 in a no contact control group. Seventeen transfer measures covered fluid intelligence, multitasking, working memory capacity, crystallized intelligence, and perceptual speed. The result was blunt: "Out of 17 ANOVAs, there were no significant Group x Session interactions," and the abstract records "no positive transfer to any of the cognitive ability tests" despite improvement on both trained tasks.
Two meta-analyses reached the same place. Monica Melby-Lervag and Charles Hulme reviewed 23 studies with 30 group comparisons in Developmental Psychology, 49(2), 270 to 291, in 2013, and found "no convincing evidence of the generalization of working memory training to other skills." Melby-Lervag, Redick, and Hulme then pooled 87 publications with 145 experimental comparisons in Perspectives on Psychological Science, 11(4), 512 to 534, in 2016. Against treated control groups, the effect on nonverbal ability was a Hedges g of .05 across 67 comparisons, with a 95 percent confidence interval running from minus .02 to .13. Against untreated controls it was .20 across 53 comparisons. The gap between those two columns is the placebo effect, measured.
What that leaves is narrow and worth stating precisely. Practice on a fluid reasoning test raises the score on that test. Training on a different task does not raise fluid reasoning in a way that shows up against an active control. The wider evidence on what does and does not move a cognitive score is on the page on improving your IQ.
10 What Score a Fluid Reasoning Test Reports
A fluid reasoning result arrives as three linked numbers, and reading it means knowing which one you are looking at. Raw score, scaled score, and index are different quantities on different scales, and a report that shows only one of them is hiding the other two.
The raw score is the count of items answered correctly on one subtest. It has no meaning on its own, because it depends entirely on which items were administered and how many there were. The scaled score converts the raw score to a standard metric against the reference group for the test taker's age, with a mean of 10 and a standard deviation of 3. The index converts the sum of scaled scores across the subtests in a domain to the familiar metric with a mean of 100 and a standard deviation of 15.
What that means in practice is that a scaled score of 13 on Matrix Reasoning is one standard deviation above the reference mean, and an FRI of 115 is the same distance on the index scale. The two are not interchangeable numbers even though they describe the same relative standing. The full mapping of scores to bands is on the IQ score scale, and how a raw performance becomes a standard score in the first place is worked through on the page on how IQ is calculated.
Every index also carries an error band. The ACIS Fluid Reasoning Index has a standard error of measurement of 2.48 points in the published normative model, which is what a confidence interval around a reported FRI is built from. An index reported without an interval is asserting a precision that no test has. The general point about measurement error is on the reliability and validity page, and the score to rank conversion is on the percentile chart.
The limits that travel with an ACIS fluid reasoning scoreACIS is a self-administered online assessment, not a clinical instrument. Its adult reference frame is documented in the technical manual. No ACIS result should be used for diagnosis, hiring decisions, accommodation requests, or high IQ society admission. What it does support is a structured reading of a person's own profile, with each index reported alongside its measurement error.
11 Reading a High or a Low Fluid Reasoning Score
An index is a position, and the useful information is usually in how it sits against the rest of the profile rather than in the number alone. Fluid reasoning correlates strongly with the other reasoning domains and weakly with speed, so what counts as an unusual pattern depends on which pair you are comparing.
In the ACIS normative model the Fluid Reasoning Index correlates .813 with Visual Spatial, .812 with Quantitative Reasoning, and .801 with Verbal Comprehension. Those are high enough that a large gap between FRI and any of the three is worth looking at. Working Memory to Processing Speed sits at .475 in the same model, so a large gap there is ordinary rather than notable. A profile is read pair by pair against the expected correlation, not by scanning for the lowest number.
Pattern
What it usually reflects
What it does not establish
FRI high, VCI similar
Broad reasoning strength picked up by both formats, which is what a high General Ability Index is built from
Any particular outcome, since the composite describes standing rather than achievement
FRI clearly above VCI
Reasoning on novel material outrunning accumulated verbal knowledge, which language background and reading history both affect
A verbal deficit, which requires an examiner and a different question to establish
FRI clearly below VCI
Knowledge outrunning novel reasoning, a pattern that is common and not by itself a concern
A reasoning impairment. An index is not a diagnosis
FRI high, PSI much lower
Very little, given the .648 g loading and 5.22 point standard error on a two subtest speed index
A processing problem, because the expected correlation between those domains is low to begin with
One FRI subtest far below the other four
Something about that specific task, from format unfamiliarity to a discontinued administration
A narrow ability deficit, which five subtests cannot separate
One further caution applies at the top of the range. Profiles at high ability levels are usually less even than profiles near the middle, and dispersion between indices at that level is expected rather than a warning. A composite still describes overall standing accurately when the profile is uneven, while the index pattern supplies direction the single figure cannot. Neither reading undermines the other, and the composite is explained on the Full Scale IQ page.
12 Taking a Fluid Reasoning Test That Reports Something
The practical question is not whether a test contains fluid items but whether it contains enough of them, in enough formats, to produce an index rather than an impression. A single matrices task is one indicator. An index needs several, and the reliability difference between two indicators and five is large enough to change how a score should be read.
ACIS reports a Fluid Reasoning Index from five subtests: Logic Grid, Figure Weights, Complex Relations, Matrix Reasoning, and Visual Number Series. Three tiers are available for adults aged 16 to 90. The Quick form at 15 dollars runs six subtests in about 45 minutes and reports a partial fluid reasoning result from two of the five. The Optimized form at 30 dollars runs 13 subtests in about 110 minutes and reports the five index profile plus the General Ability Index and the reduced verbal composite. The Full Scale form at 50 dollars runs all 20 subtests in about 175 minutes and reports the Full Scale IQ with all six primary indices and every composite.
If the question is specifically about fluid reasoning, the Optimized form is the shortest route to a complete index rather than a partial one. If the question is about where fluid reasoning sits relative to everything else, the Full Scale form is the one that reports all six domains and the composites built from them. Payment is one time with no subscription, paid access lasts 30 days, the free trial lasts seven days, and there is a five day guarantee.
What the free trial covers is worth stating exactly, because it does not cover this domain. It runs five subtests focused on working memory and processing speed, shows scaled scores and a provisional Full Scale estimate, and locks the confirmed Full Scale IQ, the six domain profile, and the percentiles. It is a way to see the administration format before paying, not a way to obtain a fluid reasoning index for nothing.
Every figure above traces to one of the following, and the entries note where a claim rests on an abstract rather than a full text.
Cattell, R. B. (1943). The Measurement of Adult Intelligence. Psychological Bulletin, 40(3), 153 to 193. The earliest published statement of the fluid and crystallized distinction, per Carroll's 1984 review of Cattell's writings. DOI 10.1037/h0059973.
Cattell, R. B. (1963). Theory of Fluid and Crystallized Intelligence: A Critical Experiment. Journal of Educational Psychology, 54(1), 1 to 22. Source of the two factor result from 277 seventh and eighth grade children. DOI 10.1037/h0046743. The full text is behind a publisher paywall, and the sample description above comes from a secondary account rather than from the article itself.
Cattell, R. B. (1971). Abilities: Their Structure, Growth, and Action. Houghton Mifflin. Revised and retitled as Intelligence: Its Structure, Growth, and Action, North-Holland, 1987.
Carroll, J. B. (1993). Human Cognitive Abilities: A Survey of Factor-Analytic Studies. Cambridge University Press, 828 pages. The publisher describes it as a reanalysis of more than 460 data sets, and no exact count is printed, so this page does not give one. Publisher page.
Schneider, W. J., and McGrew, K. S. (2018). The Cattell-Horn-Carroll Theory of Cognitive Abilities. In D. P. Flanagan and E. M. McDonough, editors, Contemporary Intellectual Assessment, fourth edition, pages 73 to 163. Guilford Press. Source of the three narrow abilities under Gf and their definitions. The earlier five ability list appears in McGrew, K. S. (2009), CHC Theory and the Human Cognitive Abilities Project, Intelligence, 37(1), 1 to 10.
Carpenter, P. A., Just, M. A., and Shell, P. (1990). What One Intelligence Test Measures. Psychological Review, 97(3), 404 to 431. Source of the five rule taxonomy, the 57 percent of error variance explained by rule count, the .91 correlation with Forbes 1964 across 2,256 British adults, the .77 Raven to Tower of Hanoi correlation across 45 students, and the goal management conclusion. Carnegie Mellon technical version at ERIC.
Gustafsson, J. E. (1984). A Unifying Model for the Structure of Intellectual Abilities. Intelligence, 8(3), 179 to 203. Source of the claim that second order Gf is identical with third order g. DOI 10.1016/0160-2896(84)90008-4.
Valentin Kvist, A., and Gustafsson, J. E. (2008). The Relation Between Fluid Intelligence and the General Factor as a Function of Cultural Background. Intelligence, 36(5), 422 to 436. Source of the .83 combined group figure and the three subgroup sample sizes. DOI 10.1016/j.intell.2007.08.004.
McGrew, K. S. (2023). Carroll's Three-Stratum Cognitive Ability Theory at 30 Years. Journal of Intelligence, 11(2), 32. Source of the counterpoint that Gf loaded .83 on g in Carroll's data, similar to comprehension knowledge, rather than approaching unity. DOI 10.3390/jintelligence11020032.
Ackerman, P. L., Beier, M. E., and Boyle, M. O. (2005). Working Memory and Intelligence: The Same or Different Constructs? Psychological Bulletin, 131(1), 30 to 60. Source of the 86 sample meta-analysis and the .479 figure. DOI 10.1037/0033-2909.131.1.30. Companion pieces in the same issue: Oberauer and colleagues at pages 61 to 65, and Beier and Ackerman's reply at pages 72 to 75.
Kane, M. J., Hambrick, D. Z., and Conway, A. R. A. (2005). Working Memory Capacity and Fluid Intelligence Are Strongly Related Constructs. Psychological Bulletin, 131(1), 66 to 71. Source of the 14 data set reanalysis across more than 3,100 subjects and the median correlation of .72. DOI 10.1037/0033-2909.131.1.66.
Jaeggi, S. M., Buschkuehl, M., Jonides, J., and Perrig, W. J. (2008). Improving Fluid Intelligence With Training on Working Memory. PNAS, 105(19), 6829 to 6833. Source of the original training claim and the sample composition. DOI 10.1073/pnas.0801268105.
Redick, T. S., and colleagues (2013). No Evidence of Intelligence Improvement After Working Memory Training. Journal of Experimental Psychology: General, 142(2), 359 to 379. Source of the three group randomized design, the 17 transfer measures, and the null result. DOI 10.1037/a0029082.
Melby-Lervag, M., and Hulme, C. (2013). Is Working Memory Training Effective? A Meta-Analytic Review. Developmental Psychology, 49(2), 270 to 291. Twenty-three studies, 30 group comparisons, no convincing evidence of generalization. DOI 10.1037/a0028228.
Melby-Lervag, M., Redick, T. S., and Hulme, C. (2016). Working Memory Training Does Not Improve Performance on Measures of Intelligence or Other Measures of Far Transfer. Perspectives on Psychological Science, 11(4), 512 to 534. Source of the 87 publication, 145 comparison pooling and the Table 1 far transfer effect sizes quoted above. Open access at PubMed Central.
ACIS technical manual. Source of every ACIS figure quoted here, including the index g loadings, the omega reliabilities, the standard errors of measurement, the subtest g loadings, and the index intercorrelations. Available at the technical manual page.
The professional framework governing all of this is explicit. The Standards for Educational and Psychological Testing, published jointly in 2014 by the American Educational Research Association, the American Psychological Association, and the National Council on Measurement in Education, require that a score interpretation be supported by validity evidence for the specific use claimed, that measurement error accompany every reported score, and that the composition and limits of the norm sample be documented. The APA standards on test use place the same duty on whoever reports a result: state what the number supports, state what it does not, and do not let a single index stand in for a person.
14 Frequently Asked Questions
What is a fluid intelligence test?
A set of problems whose rules are present in the item and absent from anyone's schooling, so the only route to an answer is to work the rule out. Matrices, series completion, quantitative balance problems, and figural analogies are the four families that do most of the measuring.
What is the difference between fluid intelligence and fluid reasoning?
Nothing substantive. Fluid intelligence is Cattell's original term and fluid reasoning is the label used in the Cattell-Horn-Carroll framework and in most current test manuals. Both refer to Gf, and index scores are usually printed under the reasoning label.
What narrow abilities sit under Gf?
Three, in the current taxonomy. Schneider and McGrew's 2018 chapter lists Induction, General Sequential Reasoning, and Quantitative Reasoning. The earlier list in McGrew's 2009 review in Intelligence also included Piagetian Reasoning and Speed of Reasoning.
When was fluid intelligence first proposed?
In 1943, in Cattell's paper The Measurement of Adult Intelligence in Psychological Bulletin. Carroll reviewed Cattell's earlier writing in 1984 and could not find the theory stated before that paper, so earlier dates are not well supported.
What did Cattell's 1963 experiment show?
That two general ability factors, not one, could be extracted from a battery combining culturally embedded and culture fair tests. The sample was 277 seventh and eighth grade children, and the result is what turned the fluid and crystallized proposal into a research program.
What does a matrix reasoning item ask you to do?
Induce the rules governing a grid of figures and apply them to a missing cell. Carpenter, Just, and Shell found in 1990 that almost every Raven problem is governed by five rule types, including constant in a row and distribution of three values.
What makes one matrix item harder than another?
Mostly the number of rules running at once. In Carpenter, Just, and Shell's 1990 analysis a regression using rule count as its only predictor accounted for 57 percent of the variance in mean error rates across the 32 problems they classified.
Can you give an example of a fluid reasoning series item?
Take 2, 6, 12, 20, 30. The gaps are 4, 6, 8, 10, so the gaps themselves form a series with a constant step of two. The next gap is 12 and the next term is 42. Nothing about the sequence has to be known in advance.
How does a figure weights or balance item work?
It states equivalences without naming them and asks you to chain them. If two circles balance three triangles and one triangle balances two squares, then four circles equal six triangles equal 12 squares. The arithmetic is trivial and the reasoning is the whole item.
Are verbal analogies fluid reasoning items?
Partly. The structure is inductive, but recognizing the relation between two words requires knowing both, so word knowledge does some of the work. That is why analogy formats appear in verbal comprehension measures too and why figural analogies are preferred for a cleaner Gf reading.
Which test is the standard measure of fluid intelligence?
Raven's Progressive Matrices, in the Advanced form for higher ability adults. It uses 48 items across a 12 item Set I and a 36 item Set II, requires no reading, and reports a score against a reference group rather than a domain profile.
What are the limits of a matrices only test?
It returns one number from one item format. It cannot say whether a low result reflects weak induction, a working memory constraint, or a strategy that stalled, and a person unusually good or bad at that specific format has no counterweight in the score.
Is fluid intelligence the same thing as g?
That is a defended position rather than a settled one. Gustafsson argued in Intelligence in 1984 that second order Gf is identical with third order g, while McGrew reported in 2023 that in Carroll's own data Gf loaded .83 on g, similar to comprehension knowledge.
Why does the ACIS Fluid Reasoning Index load highest on g?
It reaches .922 in the published technical manual, above Visual Spatial at .906 and Verbal Comprehension at .864. Both broader composites load higher still, with the General Ability Index at .954, which is what you expect if g is best estimated from breadth.
How strongly is working memory related to fluid reasoning?
That is an open dispute. Ackerman, Beier, and Boyle reported a true score correlation of .479 across 86 samples in Psychological Bulletin in 2005, and Kane, Hambrick, and Conway replied in the same issue with a median latent correlation of .72 across 14 data sets.
Why do those two numbers differ so much?
They estimate different things. A correlation between observed task scores carries each task's measurement error and task specific variance, which pulls it down. A correlation between latent factors defined by several tasks each removes that error.
Can fluid intelligence be trained?
Not by training a different task, on the current evidence. Redick and colleagues ran a randomized study with an active placebo in 2013 and reported no significant group by session interaction across 17 transfer measures despite improvement on the trained tasks themselves.
What did the meta-analyses of working memory training find?
Melby-Lervag, Redick, and Hulme pooled 87 publications with 145 comparisons in 2016. Against treated controls the effect on nonverbal ability was a Hedges g of .05 across 67 comparisons, with a confidence interval crossing zero. Against untreated controls it was .20.
What score does an ACIS fluid reasoning test report?
Scaled scores on each subtest with a mean of 10 and a standard deviation of 3, and a Fluid Reasoning Index with a mean of 100 and a standard deviation of 15. The index has an omega of .9727 and a standard error of measurement of 2.48 points.
Which ACIS tier gives a complete fluid reasoning index?
The Optimized form at 30 dollars or the Full Scale form at 50 dollars. The Quick form includes only two of the five fluid subtests and reports a partial result, and the free trial covers working memory and processing speed rather than fluid reasoning.
Is a low fluid reasoning index a diagnosis?
No. ACIS is a self-administered online assessment, not a clinical instrument, and an index describes standing against a reference frame rather than identifying a condition. Diagnosis requires a licensed examiner, a clinical battery, and a different question entirely.
Take the assessment
You get a profile, not a number
ACIS measures six CHC domains across 20 subtests and reports each one with its own normed score and confidence interval, so you can see where you are strong and where you are not.