Collective intelligence: is there a group IQ, or is it the members' IQ?
Collective intelligence is the idea that a group has a general ability of its own, a c factor, the way a person has g. A 2010 paper in Science reported one, and later work has both supported and contradicted it. This page sets out the figures each side prints, what predicts group performance, and why none of it is a statement about any one person's IQ.
The 2010 Science paper that proposed a c factor studied 40 groups of three people in its first study and 152 groups of two to five in its second, 699 participants in all.
0 The short answer
A c factor is the claim that a group has a general ability of its own, the way a person has g, and the evidence for it is real but contested. In 2010, Woolley and colleagues found that one factor carried more than 43 percent of the variance in group task scores and was not strongly correlated with the average or maximum IQ of the members. Bates and Gupta then reported that member IQ accounted for around 80 percent of the group differences in their three studies, and a 2021 meta-analysis of 22 studies found the factor again, with a published correction setting its average variance extracted at 19.6 percent. Which reading wins depends on the tasks, the interaction between members and the factor model, and none of it is a number about any one person's IQ.
699
People in the two 2010 Science studies that proposed the c factor, working in groups of two to five.
80 percent
Share of group IQ differences that individual IQ accounted for in the three studies of Bates and Gupta (2017).
0.40
Average correlation between a meta-analytic c score and performance on left out tasks in the 22 studies of Riedl and colleagues (2021).
1 What Is Collective Intelligence, and What Would a c Factor Be?
Collective intelligence, in the sense this page examines, is a measurement claim and not a slogan: that groups differ in a general ability to perform many kinds of tasks, and that one score can capture that ability. Woolley and colleagues defined a group's collective intelligence, c, as the general ability of the group to perform a wide variety of tasks, and described it as a property of the group itself and not only of the individuals in it, in their 2010 report in Science. The definition is empirical. If groups that do well on one task also tend to do well on very different tasks, and this holds across many tasks, a general factor can be extracted from the group scores. That is the operation Spearman performed on individuals' school marks in his 1904 paper on general intelligence. The first factor is called g when the cases are people and c when the cases are groups.
Three things have to hold before c earns the comparison with g, and they organize the rest of this page. First, the first factor has to carry a substantial share of the variance in group scores. Woolley and colleagues state that in individual intelligence batteries it generally accounts for 30 to 50 percent. Second, it has to carry much more than the second factor, and the caption of their scree plot describes a first factor accounting for more than twice as much variance as the later ones. Third, it has to predict group performance on other tasks better than a simpler explanation does, and the simplest rival explanation is that smart groups are made of smart people. The authors framed their own goal as testing whether the group score has predictive power beyond what knowing the individual members' abilities provides. The page on what an IQ test measures covers the individual side of the same questions.
The claim is narrower than the phrase suggests. It is not the wisdom of crowds, the finding that the pooled estimates of many independent people can be more accurate than most individuals. Lorenz and colleagues showed in an experiment with 144 participants that even mild social influence can undermine that effect in simple estimation tasks, because information about other people's answers narrowed the diversity of opinion without improving the group's collective error (PNAS, 2011). In the c factor paradigm the members work together, and Rowe, Hattie and Hester note in their 2021 meta-analysis that this differs from crowd IQ, wisdom of crowd and Delphi methods, which typically establish rules for how members divide labor, communicate and combine ideas. Whether a group of people talking, editing and deciding together has a stable general ability is a different question from whether an average of separate guesses is accurate.
Three neighboring questions belong to other pages. The individual g factor, with its own loadings and its link to the Cattell-Horn-Carroll model, is explained on the page about the g factor. The ability to read other people's emotions as a trait of one person is treated on the page about emotional intelligence and IQ. Teams of humans with artificial agents are a separate literature that this page does not cover, and the page on AI and IQ addresses what cognitive studies of language models can and cannot say. The question here is narrower: when researchers give groups batteries of tasks and analyze the group scores the way psychometricians analyze individual scores, what do they find, and how much of it is just the members' own ability?
2 How Is a Factor Extracted, and Why Is a First Factor Not Yet a g?
A first factor appears whenever tasks correlate positively, so its existence is the easy part; the contested parts are how much variance it carries, whether it survives a change of method, and what it predicts beyond the members. Spearman's observation, as Woolley and colleagues restate it, is that people who do well on one mental task tend to do well on most others, so the average correlation among scores on a diverse set of tasks is positive. A factor analysis of such scores then yields a first factor, and in individual intelligence tests that factor generally accounts for 30 to 50 percent of the variance, with later factors accounting for substantially less. The page on the g factor covers the individual version. Four statistics carry most of the debate about c, and each can be read without mathematics beyond a proportion.
Variance share. The proportion of the total variance in the task scores that the first factor takes. It depends on the extraction method: Woolley and colleagues report a first factor with an initial eigenvalue above 43 percent of the variance in Study 1, and also report that the same data yield 31 percent in Study 1 and 35 percent in Study 2 after principal axis extraction.
Loadings. The correlation between each task and the factor. A squared loading is the share of that task's variance the factor accounts for, so a loading of 0.50 means a quarter.
Average variance extracted. In its usual definition, the mean of the squared loadings of a one factor model. It answers how much of the typical task is common variance and how much is specific to the task.
Out of sample prediction. Whether a score computed from some tasks predicts performance on a task that was left out. This is the evidence that a factor is more than a summary of the tasks used to build it.
The unit of analysis is the group, which makes the samples small. The first study in Science rested on 40 groups, so each task correlation in its first table is based on 40 cases. Rowe, Hattie and Hester estimated that only 21.1 percent of the studies they reviewed had the statistical power to detect an association of the observed size, which they put as around 80 percent of studies lacking enough power. People are also nested in groups, and Credé and Howardson, whose reanalysis is described below, argued that this nesting can inflate whatever covariation exists.
A factor is also a summary, not a cause. Rowe and colleagues argue, citing the factor analysis literature, that a factor model can have mathematically equivalent solutions, so external evidence and not just statistical fit is needed to identify a factor convincingly, and they argue that for g the external evidence is criterion validity beyond the test battery itself. That is why the 2010 paper and every replication paid attention to a criterion task, and why the question of whether member IQ explains the same variance is not a side issue. The table sets out the questions that decide the claim.
Question
Why it decides the claim
Where it is examined below
Do group scores on different tasks correlate positively?
Without a positive manifold there is nothing to extract.
The 2010 studies and the 22 study meta-analysis
Does one factor carry a large share, and at least twice the next?
A factor that explains little is a weak summary.
The 2010 studies, Credé and Howardson, the 2021 correction
Does the factor appear in other kinds of groups and tasks?
A factor tied to one task set is a property of the tasks.
Online groups, task type
Does it predict a later task beyond what member IQ predicts?
This separates c from "smart members make a smart group".
Bates and Gupta, Rowe and colleagues
Is the factor stable across analytic choices?
A factor that disappears under another model is fragile.
Credé and Howardson, the two factor model
3 What Did Woolley and Colleagues Find in 2010?
The 2010 Science report found, in both of its studies, one dominant factor in group task scores and a c score that predicted a later criterion task far better than the average or maximum IQ of the members. In two studies with 699 people, working in groups of two to five, the authors report converging evidence of a general collective intelligence factor. In Study 1, 40 three person groups, formed by random assignment, worked for up to five hours on diverse simple tasks plus a more complex criterion task, which was playing checkers against a standardized computer opponent. The tasks were sampled from all quadrants of the McGrath task circumplex, a taxonomy of group tasks, and included solving visual puzzles, brainstorming, making collective moral judgments and negotiating over limited resources. Members' individual intelligence was measured at the start of each session. Study 2 used 152 groups of two to five members, a broader set of tasks, an alternative measure of individual intelligence and an architectural design task as the criterion.
The correlations among group tasks were positive, with an average inter task correlation of 0.28 in Study 1. The first factor had an initial eigenvalue above 43 percent of the variance in Study 1, against 18 percent for the second, and 44 percent in Study 2, against 20 percent. In a subset of Study 2 groups the authors added five tasks for a total of ten and report results consistent with a single factor. A confirmatory one factor model fit both studies, and the authors report that the factor structures of the two studies were invariant.
Measure
Study 1
Study 2
Groups
40, three members each
152, two to five members
First factor, initial eigenvalue share of variance
More than 43 percent
44 percent
Second factor
18 percent
20 percent
First factor after principal axis extraction
31 percent
35 percent
Criterion task
Checkers against a computer
Architectural design
Regression weight of c on the criterion
0.51
0.36
Regression weight of average member intelligence
0.08, not significant
0.05, not significant
Regression weight of maximum member intelligence
0.01, not significant
0.12, not significant
The criterion results carry the claim of added value. In Study 1 the c score correlated 0.52 with checkers performance, while average and maximum member intelligence correlated 0.18 and 0.13 with it, neither significant. In Study 2, 63 individuals who completed the design task alone showed that individual intelligence did predict it (r = 0.33), yet the average intelligence of group members did not predict group performance on the same task (r = 0.18, not significant). When both c and average individual intelligence were entered together, c was a significant predictor in both studies and member intelligence was not. This is the basis of the statement in the abstract that c is not strongly correlated with the average or maximum individual intelligence of group members. The authors also report that group cohesion, motivation and satisfaction, which one might expect to matter, did not predict c.
Three features of this evidence matter for what follows. The sample is small at the level that counts, 192 groups in all. The Study 1 criterion task was played at the end of the same session as the battery, so c predicted a later task and not a later outcome in a workplace. And the authors presented their own limits: the paper ends with open questions about whether a short collective intelligence test could predict a sales team's or a management team's long term effectiveness, and whether c can be raised more easily than individual intelligence. The next sections follow what other laboratories found when they asked the same questions.
4 Does the Factor Appear in Online Groups? Two Studies That Disagree
Two studies of groups that communicate through a computer reached opposite conclusions about whether a c factor appears, which suggests that the factor depends on the tasks and the setting. Engel, Woolley, Jing, Chabris and Malone ran the 2010 design with a new task battery and manipulated whether members were in the same room. In their 2014 paper in PLOS ONE, the first factor explained 49 percent of the variance in face to face groups and 41 percent in online groups, and the second factor explained less than half as much in each case. The online groups, 36 of them against 32 face to face, communicated only by text and never saw each other.
The same paper tested the 2010 finding about social sensitivity. The members' average score on the Reading the Mind in the Eyes test, which requires participants to read the mental states of others from looking at their eyes, correlated 0.53 with c in face to face groups and 0.55 in online groups. The authors read this as evidence that the test taps a broader ability to reason about other people's mental states and not only the recognition of facial expressions, since the online groups had no faces to read. That is an interpretation of the correlation, not a result of the experiment.
Barlow and Dennis studied groups using computer mediated communication and did not find a factor. Their 2016 paper in the Journal of Management Information Systems reports that a collective intelligence factor did not emerge among those groups, concludes that collective intelligence manifests itself differently depending on context, and calls for research on the boundary conditions of the construct. Rowe and colleagues examined the correlation matrices in the literature and found that the extremely low and negative correlations among group tasks were mostly attributable to Barlow's studies. In one of them, a profit maximization task chosen as a more complex criterion correlated 0.02, 0.01 and 0.15 with the brainstorming, college admissions and shopping trip tasks. Rowe and colleagues read this as a problem with the task set, noting that evidence for the reliability and validity of those tasks as ability measures is lacking. A reader can hold both findings at once: online groups produced a factor in one battery and not in another, so the answer turns on which tasks were chosen and how much they let members interact.
5 What Did the Critiques and Failed Replications Find?
Independent analyses reported that the c factor is weak, depends on the tasks, or reduces to member IQ, and each of the four main challenges used a different method. Credé and Howardson reexamined six previously published samples in the Journal of Applied Psychology in 2017. They report that the general factor explains only little variance in performance on many group tasks and that two statistical artifacts, the apparent presence of low effort responding and the nested nature of the data, may have inflated the little covariation that exists. Their conclusion is that there is insufficient support for the existence of a collective intelligence construct. Rowe and colleagues describe their argument as resting on simulations in which partitioning the individually nested data removed most of the common variance at the group level. Woolley, Kim and Malone posted a reply to the critique on the Social Science Research Network in 2018, which this page lists by title only because its content was not read.
Bates and Gupta tested the 2010 model directly. Their 2017 paper in Intelligence reports three studies with 312 people. Contrary to the prediction, individual IQ accounted for around 80 percent of group IQ differences, the hypotheses that group IQ rises with the number of women and with turn taking were not supported, and performance on the Reading the Mind in the Eyes test was associated with individual IQ and, in one study, with group IQ factor scores. A well fitting structural model combining studies 2 and 3 indicated that the eyes test exerted no influence on the latent group factor, and in that model individual IQ determined 100 percent of the latent group IQ. Rowe and colleagues point out that the task correlations in these three studies were among the strongest in the literature and involved items known to load heavily on psychometric g, for example a missing letters task that correlated 0.80 with word fluency and 0.64 with Raven's matrices. That fits a simple reading: when the group tasks are themselves g loaded ability tasks, the groups with the most able members do best, and a group factor then largely reproduces member ability.
Rowe, Hattie and Hester pooled what existed in a 2021 meta-analysis. Across 19 samples, only 7 were independent of Woolley and affiliated coauthors, and they report that all attempts to reproduce the c factor under independent authorship either partially or completely failed, while all Woolley affiliated attempts succeeded. Across eight independent samples (857 groups) the c factor correlated 0.26 (95 percent confidence interval 0.10 to 0.40) with nine criterion tasks. Only four studies, five independent samples and 366 groups, controlled for member IQ, and in that subset average member IQ had little to no correlation with group performance (r = 0.06, 95 percent interval -0.08 to 0.20). They also report that about 80 percent of studies lacked the power to detect the observed associations. Their conclusion is balanced, not dismissive: some findings are consistent with a general factor of group performance that relates positively to performance, alternative explanations cannot be dismissed, and in their opinion the case against the c factor has not been firmly established. They advise against embracing it until it is independently replicated and shown to add validity beyond g.
A later experiment pushed in a different direction. In PLOS ONE in 2024, Rowe, Hattie and Munro gave 85 university students, allocated to 29 groups of two to five, a series of cognitive tasks, and found that a two factor model fit better than the single factor model, with the two factors modeled on Cattell's fluid and crystallized intelligence. Both individual and collective intelligence measures correlated moderately with group assignment grades (r from 0.40 to 0.47). The sample is small and the finding is one study, but it matters conceptually because it suggests that the group level structure may resemble the multi factor structure of individual intelligence, which the page on fluid and crystallized intelligence describes, and not a single g.
Source
Design and sample
What it reported
Barlow and Dennis, 2016
Groups using computer mediated communication
No collective intelligence factor emerged
Credé and Howardson, 2017
Reanalysis of six published samples
General factor explains little variance on many tasks; insufficient support
Bates and Gupta, 2017
Three studies, 312 people
Individual IQ accounted for around 80 percent of group IQ differences; women and turn taking not supported
Rowe, Hattie and Hester, 2021
Meta-analysis, 19 samples, 7 independent of the original authors
c to criterion r = 0.26; member IQ to performance r = 0.06 in the controlled subset; most studies underpowered
Rowe, Hattie and Munro, 2024
85 students in 29 groups
Two factor model fit better than one factor
6 What Did the 2021 Meta-Analysis of 22 Studies Find, and What Did Its Correction Change?
The largest test of the c factor, a 2021 meta-analysis of 22 studies, supported a single factor and out of sample prediction, but its headline variance figure was corrected in 2022 from 44 percent to 19.6 percent. Riedl, Kim, Gupta, Malone and Woolley analyzed group performance data from 22 studies including 5,279 individuals in 1,356 groups, in PNAS in 2021. They conclude that a robust collective intelligence factor characterizes a group's ability to work together across a diverse set of tasks, that it is predicted by the proportion of women in the group, mediated by average social perceptiveness, and that it predicts performance on out of sample criterion tasks. They also report that group collaboration process is more important in predicting c than the skill of individual members.
The figures come from the paper itself. Group scores across the tasks correlated 0.27 on average, with a range from 0.12 to 0.50. A one factor model had standardized loadings from 0.27 to 0.52, all significant. To test prediction, the authors computed a c score from seven of the eight tasks and used it to predict the task left out, repeating this eight times, and the average correlation was 0.40 (0.26 to 0.53). The importance of process versus skill varies by task: more than 51 percent of the explained variation in Sudoku performance was due to individual member skill, while for unscrambling words, 55 percent of the variation was due to group processes. They also report negative effects of high age diversity on c.
A correction published in PNAS in 2022 changed one headline figure. The authors wrote that, due to a mistake in a software package, the average variance extracted for the factor analyses had been reported incorrectly. The reported average variance extracted of 44 percent was replaced with 19.6 percent, some model fit statistics in the supplement were revised, and the authors state that the corrections do not change the major conclusions of the paper. The factor loadings, 0.27 to 0.52, stand unchanged. Squared, those loadings run from about 0.07 to about 0.27 (our arithmetic), which is consistent with a typical task sharing about a fifth of its variance with the factor and not with 44 percent.
That corrected 19.6 percent is not a like for like comparison with the 43 and 44 percent of 2010. The 2010 figures are the share of variance taken by the first eigenvalue, and the authors' own principal axis estimates for the same data were 31 and 35 percent, while the 2021 figure is the average variance extracted of a confirmatory model fitted to pooled correlations. Anyone who sets one beside the other is comparing different statistics, so this page does not subtract them. What the corrected figure does show is that the factor, though present, accounts for a modest part of the variance in any given task, which is close to the Credé and Howardson observation that the factor carries little variance on many tasks, although the meta-analysis authors read their results as support for a robust single factor.
Two cautions apply when reading the meta-analysis. First, two of its five authors, Woolley and Malone, are authors of the 2010 paper, and Rowe and colleagues counted only 7 of 19 samples as independent of the original authors, which is why authorship matters when reading this literature. Second, the skill measure the authors describe in the open preprint of the paper is an estimate of each member's skill on the tasks performed, not a score from an IQ test, so the meta-analysis is not a head to head test of c against member IQ on the same criterion. That comparison is the one the next section turns to.
7 Does Member Ability Explain Group Performance? Average and Maximum IQ
Member ability is positively related to group performance in nearly every source, and the sources differ on how much it explains, partly because they report different quantities. The 2010 paper pooled both studies and found that average member intelligence correlated 0.15 with c and the intelligence of the highest scoring member 0.19, which the authors called moderate. Rowe and colleagues report that three earlier meta-analyses of general intelligence and group performance (Bell in 2007, Devine and Philips in 2001 and Stewart in 2006), combined, give a sample weighted correlation of 0.28 between average member IQ and group performance, with a 95 percent interval of 0.25 to 0.30. That 0.28 is not the same quantity as the 0.15, because one is a correlation with performance on actual tasks and the other is a correlation with a derived group factor.
Source
Quantity reported
Value
Woolley and colleagues, 2010
Average and maximum member intelligence with c (pooled)
0.15 and 0.19
Three earlier meta-analyses, as summarized by Rowe and colleagues
Average member IQ with group performance
0.28 (95 percent interval 0.25 to 0.30)
Rowe, Hattie and Hester, 2021
Average member IQ with c (five samples, 345 groups)
0.19; 0.32 in independent studies, 0.10 in Woolley affiliated studies
Rowe, Hattie and Hester, 2021
Average member IQ with criterion tasks (five independent samples, 366 groups)
0.06 (interval -0.08 to 0.20)
Bates and Gupta, 2017
Share of group IQ differences accounted for by individual IQ
Around 80 percent; 100 percent of the latent factor in the combined model
Blanchard and colleagues, 2025
Individual intelligence as a predictor of dyadic c, standardized weight
0.33
The most recent entry comes from a study built to address task type. Blanchard, Aidman, Stankov and Kleitman measured collective intelligence in 105 undergraduate dyads using three group tasks aligned with broad abilities of the Cattell-Horn-Carroll model, and assessed individual intelligence with Raven's Advanced Progressive Matrices. In their 2025 paper in Cognitive Research: Principles and Implications, individual intelligence (standardized weight 0.33) and individual confidence (0.32) were the strongest predictors of dyadic c for well structured tasks, together adding 25 percent of the variance after social sensitivity, gender composition, working memory and personality had been entered. The authors suggest that the earlier emphasis on social factors may depend on the type of tasks groups complete, and they underline the importance of task selection. The page on the CHC model explains the broad abilities their tasks were aligned with.
Why do sources that look at the same question land so far apart? Four reasons recur in the papers themselves or follow from ordinary psychometrics, and the last two are our reading more than an author's claim.
Task type. In the meta-analysis of Riedl and colleagues, more than 51 percent of the explained variation in Sudoku was due to member skill, and for unscrambling words, 55 percent was due to process. A battery of closed ended tasks with one correct answer, which any capable member can find, will make member ability matter more than a battery of brainstorming and typing.
Selection of the tasks. Bates and Gupta used tasks that load heavily on g, and Rowe and colleagues observed that this produced some of the highest task correlations in the literature.
How member ability was measured. Rowe and colleagues' table notes that one 2010 study used 18 of the 36 items of Raven's Advanced Progressive Matrices and that the Wonderlic Personnel Test was used in most of the others, and that ceiling effects may have partly concealed the effect of member IQ in at least one study. A short or ceilinged measure of an ability correlates less with anything, which is a general psychometric point.
Small samples at the level of the group. With roughly 25 to 150 groups in each of the individual studies tabulated here, differences of a few hundredths in a correlation cannot be told from noise, which is what an underpowered literature means in practice.
The conclusion this supports is modest. Nobody in these sources reports that member ability is irrelevant, and nobody reports that it is the only thing that matters. The sources differ on its share, and the share changes with the tasks.
8 Does Social Perceptiveness Predict Group Performance?
Social perceptiveness, measured by one test of reading mental states from the eyes, is the predictor with the most consistent support in these sources, and also the one most entangled with individual ability. In 2010 the average score of group members on the Reading the Mind in the Eyes test correlated 0.26 with c (P = 0.002). Engel and colleagues found 0.53 in face to face groups and 0.55 in online groups, and Riedl and colleagues report that average social perceptiveness mediates the effect of the proportion of women on c. Blanchard and colleagues replicated the relationship in dyads, with a standardized weight of 0.22. In a 2024 paper in Perspectives on Psychological Science, Woolley and Gupta describe a Transaction Systems Model of Collective Intelligence in which transactive memory, attention and reasoning systems develop and adapt together to support collective intelligence. That is a proposed account of mechanism, and the abstract read for this page does not report a test of it against the member ability account.
Two refinements matter. Meslec, Aggarwal and Curseu tested which members' scores count, in a 2016 paper in Frontiers in Psychology. They report that collectively intelligent groups are those in which the least socially sensitive member has a rather high score, so that sensitive members cannot compensate for the lack of it in others. And Bates and Gupta found that eyes test performance was associated with individual IQ, which means a correlation of the test with group performance can partly reflect ability, and that in their combined model it had no influence on the latent group factor. Whether this kind of perceptiveness is a trait distinct from cognitive ability is the subject of the page on emotional intelligence and IQ. For the c factor, the safe reading is that a mind reading measure predicts group performance in several samples, and that its independence from member IQ is not settled.
9 Do Equal Turn Taking and More Women Make Groups Smarter?
Equal turn taking and the share of women were two of the three correlates reported in 2010, and they replicated less consistently than social sensitivity, so a page that states them as established is ahead of the evidence. In 2010, c correlated negatively with the variance in the number of speaking turns, measured with sociometric badges worn by a subset of groups (r = -0.41, P = 0.01), so groups dominated by a few speakers scored lower. It also correlated positively with the proportion of women in the group (r = 0.23, P = 0.007). The authors say that this second result appears to be largely mediated by social sensitivity, because the women in their sample scored better on the eyes test than the men. In a regression where all three predictors were available, only social sensitivity reached significance (weight 0.33, P = 0.05).
The later record is mixed.
Predictor
Woolley and colleagues, 2010
Later work that agrees
Later work that does not
Social sensitivity
r = 0.26 with c
Engel 2014 (0.53 and 0.55); Riedl 2021; Blanchard 2025 (weight 0.22); Meslec 2016 (lowest member matters)
Bates and Gupta 2017: no influence on the latent factor in the combined model
Equal turn taking
r = -0.41 with turn variance
None found in the sources read for this page
Bates and Gupta 2017: not supported; Blanchard 2025: no effect (weight 0.08)
Proportion of women
r = 0.23, largely mediated by social sensitivity
Riedl 2021: positive, mediated by social perceptiveness
Bates and Gupta 2017: not supported; Blanchard 2025: not supported
Blanchard and colleagues also report what they call a strategic dominance pattern: some dyads performed better together than alone by letting the more competent member dominate discussion and decisions, which, in their words, challenges the earlier claim about equal turn taking. They note that Rowe and colleagues' 2024 study found neither social sensitivity nor turn taking related to the collective intelligence factors. In Blanchard's dyads, the proportion of women did not help either: all female dyads scored lower than all male dyads (standardized weight -0.54, P = 0.01), while mixed dyads did not differ from all male dyads. That is a result about undergraduate pairs on well structured tasks, and it supports a sweeping claim about women no more than the 2010 correlation did.
Two cautions about reading such results apply to every row. A correlation between a group's makeup and its score is a statement about groups of people, and a difference between group averages on one test says nothing about any individual woman or man. And "mediated" is statistical: it means the association runs through the average eyes test score in the data, not that a causal mechanism has been shown.
10 Where Do the Two Sides Agree, and Where Do They Not?
The two sides agree more than the headlines suggest: group performance across tasks correlates positively, member ability matters and the tasks chosen change the answer, and they disagree about how large the group factor is and what it adds. The first point of agreement is that group scores on different tasks correlate positively, at about 0.27 on average in the 22 studies of Riedl and colleagues and 0.28 in the first 2010 study. The second is that member ability is positively related to group performance in most of the sources above, though near zero in one subset. The third is that the findings vary with the type of task, which Riedl and colleagues demonstrate within their own data and Blanchard and colleagues argue as a reason for the divergence.
The disagreements are also specific. Credé and Howardson hold that the general factor explains little variance and that artifacts may inflate it, whereas Riedl and colleagues report a robust factor whose corrected average variance extracted is 19.6 percent. Bates and Gupta report that member IQ determines the latent group factor, whereas Woolley and colleagues report that c predicted criterion tasks where member intelligence did not. Rowe and colleagues report that independent replications fared worse than affiliated ones, and Riedl and colleagues point to methodological choices, such as restricting how many members can record answers, which they say would obscure group capability on process dependent tasks. Rowe, Hattie and Munro suggest two factors where the others find one.
One possible reconciliation, which is ours and has not been tested head to head in the sources read, is that both are partly right for different tasks: closed ended tasks with a single correct answer reward the ablest member, while generative and executing tasks reward coordination. The Riedl split between Sudoku and unscrambling words and the dyad findings of Blanchard are consistent with that, and it would also explain why selecting tasks that load on g produced groups whose group IQ was individual IQ. Until a study varies task type and member ability together at an adequate sample size, a reader should treat the c factor as a promising and partly replicated description of how groups perform on a given set of tasks, not as a settled trait of groups.
11 What Does This Say About Your Own Team, and Your Own IQ?
Nothing in these studies converts a group's score into a person's IQ, or one person's IQ into a verdict on a team, because every result is a statistic about groups under particular tasks. The cases in all of these analyses are groups of strangers, friends or classmates working on research tasks for a few hours. A c score describes how a set of groups ranked on a battery, in the same way that a school's average score ranks schools and not students. The 22 studies of Riedl and colleagues do not describe a standing team at work, and the 2010 criterion tasks were checkers and an architectural design exercise, not a quarter of business results.
The individual side has its own literature, which answers a different question. Whether general mental ability predicts job performance is treated in the page on IQ and job performance, and what an employer's test is and is not appears in the page on pre-employment cognitive tests. The page on the Pygmalion effect covers a separate route by which others' beliefs can shape performance, and the page on what intelligence is places individual ability among the definitions researchers use. Rowe and colleagues' own practical advice, after weighing the c factor evidence, is that researchers and practitioners continue to measure and account for intelligence in groups with individual IQ tests validated around psychometric g.
ACIS measures individuals and does not measure groups. Its report gives a Full Scale IQ and six index scores on the standard scale, a percentile and a 95 percent confidence interval, as described on the page about the six cognitive domains. It is an online, unsupervised, English only assessment. It is not a clinical or diagnostic instrument and it is not for hiring, school accommodations or admission to high IQ societies. A person's report says nothing about how well that person collaborates, and a team's results on a task battery would not be an ACIS result.
12 How Should a Group Score Be Read? What the Standards Say
The Standards for Educational and Psychological Testing treat validity as a property of an interpretation for a proposed use, so evidence that supports reading c as a group trait would not support reading anything about one member, and the reverse. The Standards for Educational and Psychological Testing, published jointly by AERA, APA and NCME in 2014, define validity as the degree to which evidence and theory support the interpretations of test scores for proposed uses, say that the interpretations and not the test are what is evaluated, and note that each intended interpretation must be validated. The full text of the Standards names two kinds of evidence that map directly onto this debate. Evidence based on internal structure asks whether the relationships among components fit the construct, which is what the factor analyses of Woolley, Credé and Howardson and Rowe address. Evidence based on relations to other variables includes convergent and discriminant evidence, which is what the comparison of c with member IQ tests.
The Standards also contain a rule on aggregation, Standard 6.12: when group level information is obtained by aggregating the results of partial tests taken by individuals, evidence of validity and reliability or precision should be reported for the level of aggregation at which results are reported. The standard concerns large scale assessment, so applying it here is our analogy, but its logic carries over: evidence about individual scores does not transfer to group scores, and a group score needs evidence of its own. The APA Guidelines for Psychological Assessment and Evaluation, approved by the APA Council of Representatives in March 2020, add in Guideline 5 that psychologists should apply psychometric principles together with the effects of external sources of variability such as context, setting, purpose and population. A c score from laboratory groups says little about that context, setting and population when the real question is about a standing team.
A practical test follows from these documents, and it is our synthesis and not a list from them. When someone claims a team has high collective intelligence, ask which tasks produced the score, how many groups were in the sample, whether member IQ was measured and controlled, whether the factor was checked on tasks that were left out, and whether the claim is about the group or about one of its members. The page on reliability and validity explains the same questions for an individual test, and the technical manual documents how ACIS reports individual scores.
Every figure on this page comes from the sources below, which were checked on October 6, 2026, and where a source was available only as an abstract the page says so. DOIs were confirmed against Crossref. The 2010 Science report was read in a copy of the printed article, the 2014 Engel paper, the 2021 Rowe meta-analysis, the 2024 Rowe paper and the 2025 Blanchard paper in their open full texts, the 2021 Riedl paper through its PubMed Central page and an open preprint, and its correction in full. The Bates and Gupta, Credé and Howardson, Barlow and Dennis, Meslec, Woolley and Gupta and Lorenz papers were read as abstracts, and statements about the first three beyond their abstracts are attributed to Rowe and colleagues. The Standards and the APA guidelines were read as documents.
Woolley, A. W., Chabris, C. F., Pentland, A., Hashmi, N., and Malone, T. W. (2010). Evidence for a Collective Intelligence Factor in the Performance of Human Groups. Science, 330(6004), 686 to 688.
Lorenz, J., Rauhut, H., Schweitzer, F., and Helbing, D. (2011). How social influence can undermine the wisdom of crowd effect. Proceedings of the National Academy of Sciences, 108(22), 9020 to 9025.
Engel, D., Woolley, A. W., Jing, L., Chabris, C. F., and Malone, T. W. (2014). Reading the Mind in the Eyes or Reading between the Lines? Theory of Mind Predicts Collective Intelligence Equally Well Online and Face-To-Face. PLOS ONE, 9(12), e115212.
Barlow, J., and Dennis, A. (2016). Not As Smart As We Think: A Study of Collective Intelligence in Virtual Groups. Journal of Management Information Systems, 33(3), 684 to 712.
Meslec, N., Aggarwal, I., and Curseu, P. (2016). The Insensitive Ruins It All: Compositional and Compilational Influences of Social Sensitivity on Collective Intelligence in Groups. Frontiers in Psychology, 7, 676.
Credé, M., and Howardson, G. (2017). The structure of group task performance: A second look at "collective intelligence": Comment on Woolley et al. (2010). Journal of Applied Psychology, 102(10), 1483 to 1492.
Bates, T. C., and Gupta, S. (2017). Smart groups of smart people: Evidence for IQ as the origin of collective intelligence in the performance of human groups. Intelligence, 60, 46 to 56.
Riedl, C., Kim, Y. J., Gupta, P., Malone, T. W., and Woolley, A. W. (2021). Quantifying collective intelligence in human groups. Proceedings of the National Academy of Sciences, 118(21), e2005737118.
Rowe, L. I., Hattie, J., and Hester, R. (2021). g versus c: comparing individual and collective intelligence across two meta-analyses. Cognitive Research: Principles and Implications, 6(1), article 26.
Rowe, L. I., Hattie, J., and Munro, J. (2024). High-performing teams: Is collective intelligence the answer? PLOS ONE, 19(8), e0307945.
Woolley, A. W., and Gupta, P. (2024). Understanding Collective Intelligence: Investigating the Role of Collective Memory, Attention, and Reasoning Processes. Perspectives on Psychological Science, 19(2), 344 to 354.
Blanchard, M. D., Aidman, E., Stankov, L., and Kleitman, S. (2025). A recipe for dyadic collective intelligence for well-structured tasks: mix equal parts cognitive ability and confidence plus a pinch of social sensitivity. Cognitive Research: Principles and Implications, 10(1), article 63.
American Educational Research Association, American Psychological Association, and National Council on Measurement in Education (2014). Standards for Educational and Psychological Testing. Washington, DC: AERA. Read October 6, 2026.
American Psychological Association, APA Task Force on Psychological Assessment and Evaluation Guidelines (2020). APA Guidelines for Psychological Assessment and Evaluation. Approved by the APA Council of Representatives in March 2020, read October 6, 2026.
14 Frequently Asked Questions
What is collective intelligence in psychology?
In psychology it usually means a group's general ability to perform many different kinds of tasks, as opposed to its skill at one task. Researchers measure it by giving groups several tasks and extracting a single factor from the group scores, the way a general factor is extracted from individuals' scores.
What is the c factor?
The c factor is the first factor extracted from a group's scores across varied tasks, proposed in 2010 as the group counterpart of g. A group with a high c score tends to do well on many kinds of tasks, and whether it is more than the members' own ability is the contested point.
Is there such a thing as group IQ?
Only as a research construct. Researchers have scored groups on task batteries, and some studies find a general factor, but others find that member IQ explains most of the differences. A group IQ is therefore a contested research score and not a settled trait of groups.
Is collective intelligence real?
Groups clearly differ in how well they perform across tasks, and that ranking is partly stable. It is unsettled whether the stability is a distinct group ability or mostly the members' abilities at work. Independent analyses have been less supportive than the original authors' studies, so the honest answer is partly.
Is collective intelligence the same as the wisdom of crowds?
No. The wisdom of crowds concerns pooling many independent estimates, which can beat most individuals. The c factor concerns groups whose members interact while working on tasks. Experimental work shows social influence can erode the crowd effect, so the two are different phenomena.
Who discovered the collective intelligence factor?
Anita Williams Woolley, Christopher Chabris, Alex Pentland, Nada Hashmi and Thomas Malone proposed it in a 2010 paper in Science. Their design borrowed Spearman's method for individual intelligence, which dates to 1904. Later researchers, including some of the same authors, tested and extended the idea.
Is a smart group made of smart people?
Partly. Member ability correlates positively with group performance in most studies, but the correlations reported range from near zero to about 0.3, and one study reported that individual IQ accounted for around 80 percent of group IQ differences. The amount depends on the tasks and the measures.
What did the 2010 Woolley study find?
Across two studies of 699 people in groups of two to five, one factor dominated group task scores, and a c score predicted later criterion tasks better than average or maximum member intelligence did. It also correlated with social sensitivity, equal speaking turns and the share of women in the group.
Why did Bates and Gupta find different results?
Their three studies used tasks that load heavily on psychometric g, so group IQ tracked member IQ closely. Rowe and colleagues noted that these studies had some of the strongest task correlations in the literature. Different tasks, different samples and small numbers of groups can each change the answer.
What did the 2021 meta-analysis conclude?
Riedl and colleagues analyzed 22 studies and concluded that a single collective intelligence factor characterizes group performance across diverse tasks, predicts tasks left out with an average correlation of 0.40, and depends more on collaboration process than on member skill. Its variance figure was later corrected.
Why was the 2021 meta-analysis corrected?
The authors reported that a mistake in a software package made the average variance extracted for their factor analyses incorrect. The figure printed as 44 percent became 19.6 percent, some supplement fit statistics were revised, and they stated that the major conclusions of the paper did not change.
What did Credé and Howardson argue?
They reanalyzed six published samples and argued that the general factor explains little variance in many group tasks, and that low effort responding and the nesting of people within groups may inflate the covariation. They concluded that support for a collective intelligence construct is insufficient.
Does equal turn taking make groups smarter?
The 2010 study found that groups with more equal speaking turns scored higher, but later work did not support it. Bates and Gupta did not find the effect, and Blanchard and colleagues found none in dyads, where letting the more competent member dominate sometimes helped performance.
Do groups with more women perform better?
The 2010 study found a modest association, largely explained by higher average scores on an eyes based social sensitivity test, and the 2021 meta-analysis reported the same mediation. Two later studies did not support it. It is a statement about groups, not about any individual.
Can an IQ test measure a team's intelligence?
An individual IQ test measures the individual, not the team. It can describe each member's abilities, which matter in many studies, but it cannot return a team score. Whether a team has its own general ability is a research question that needs group tasks, not individual scores.
Does the highest IQ in a group decide how the group performs?
Not by itself. In the 2010 studies the highest scoring member's intelligence correlated 0.19 with c and did not significantly predict the criterion tasks. Other studies give member ability more weight, especially on tasks with one correct answer, so it depends on what the group is doing.
Does social sensitivity matter more than IQ for teams?
The evidence does not support ranking them. Social sensitivity predicted c in several studies, but Bates and Gupta found the eyes test related to individual IQ, and Blanchard and colleagues found individual intelligence and confidence the strongest predictors for well structured dyad tasks.
Can collective intelligence be used to hire or choose teams?
This evidence does not support that. The studies used laboratory groups on short tasks, the factor itself is disputed, and the criterion tasks were not workplace outcomes. A selection use needs its own validation for that purpose, as the Standards for Educational and Psychological Testing describe.
Does a group's collective intelligence say anything about my own IQ?
No. A c score is a statistic about how groups ranked on tasks, and it cannot be converted into an individual score or the reverse. Your IQ describes your own performance against a norm group and does not tell how any team you join will perform.
Does ACIS measure group or collective intelligence?
No. ACIS administers an individual assessment and reports a Full Scale IQ and six index scores with percentiles and a 95 percent confidence interval. It is online and unsupervised, not clinical or diagnostic, and not for hiring, school accommodations or admission to high IQ societies.
How should I read a claim that a team has high collective intelligence?
Ask which tasks produced the score, how many groups were studied, whether member IQ was measured and controlled, and whether the score predicted tasks that were left out. If the claim rests on one short battery or a survey, treat it as a description of that battery, not a trait.
Take the assessment
You get a profile, not a number
ACIS measures six CHC domains across 20 subtests and reports each one with its own normed score and confidence interval, so you can see where you are strong and where you are not.