Where intelligence sits in the brain: what imaging studies replicate
Intelligence is not located in one region of the brain. Imaging studies link higher scores to a distributed network that includes the frontal and parietal cortex, to the white matter fibers between them, and to how efficiently the network communicates. The effects are modest and many early findings failed to replicate. This page reports the effect sizes each source prints and separates what holds from what does not.
A single drawn brain sits in one place, whereas the P-FIT review summarized here drew on 37 imaging studies and located the signal in a network across frontal, parietal, temporal and occipital regions.
0 The short answer
Intelligence is not in one place: neuroimaging links higher scores to a distributed network centered on the frontal and parietal cortex, plus the white matter between them, and the associations are modest. Total brain volume correlates with IQ at about r = .24 in a meta-analysis of 88 studies, and the P-FIT model synthesized 37 imaging studies into a parieto-frontal network. The limit is replication: in about 50,000 people the median brain-wide association with cognition and mental health measures was tiny (r = 0.01 in the main sample), so small studies overstate effects. These are group correlations, not a way to read one person's IQ from a scan.
37
Imaging studies that Jung and Haier synthesized in 2007 into the P-FIT model of a distributed parieto-frontal network.
r = .24
Average association between brain volume and IQ in the 2015 meta-analysis of 88 studies, which explains about 6 percent of variance.
0.01
Median univariate brain-wide association (absolute r) in the 3,928 person main sample of Marek and colleagues, from a total of about 50,000 people.
1 Why Is It Hard to Say Where Intelligence Is in the Brain?
The question has no single answer because intelligence research correlates an estimated general ability with brain measurements, and the size of every finding depends on how well each side was measured. Richard Haier and Rex Jung, in their chapter on the parieto-frontal integration theory in the fourth edition of Contemporary Intellectual Assessment (Guilford Press, 2018), make a point that shapes everything below: the general factor g is estimated, not measured. No instrument reads intelligence as a quantity the way a scale reads weight. A score has meaning only in relation to other people, which is why the chapter says test scores are best interpreted as percentiles. The page on how g is extracted from a battery explains the estimation. What matters here is the consequence: when one study uses a single test and another uses ten, they are not measuring the same thing, and a weaker brain correlation may only reflect a weaker measure.
The brain side is just as varied. Researchers use several techniques, and each answers a different question.
Technique
What it measures
Typical design
What it can support
Positron emission tomography (PET)
Glucose use by cortical regions during a task
Small samples, scanned while solving problems
Which regions work harder or less hard during a task, in the people scanned
Structural MRI
Volume or thickness of gray matter regions, and total brain volume
Large samples possible, one scan per person
Whether people with larger values tend to score higher
Diffusion MRI
Integrity of white matter tracts, such as fractional anisotropy
Medium to large samples
Whether better connected fibers tend to go with higher scores
Functional MRI during tasks
Regional activation while people reason
Small to medium samples
Which regions are engaged by hard reasoning problems
Resting-state functional MRI
Statistical coupling between regions while a person lies still
Large public datasets
Whether connection patterns predict scores in new people
Lesion mapping
Which damaged regions go with lower scores in patients
Patients with focal brain damage
Which regions are necessary for performance, in the patients studied
Read the right-hand column as a limit. Almost all of these designs observe each person at one point in time, so they show that a brain measure and a score go together, not that one produces the other. Lesion mapping comes closest to a claim about necessity, because damage that removes a region and lowers performance is harder to explain by chance, but even it studies a special group of patients. Throughout this page the verbs follow the design: "is associated with," "predicts in new samples," "goes with." A causal verb appears only where a design earns it.
Two related questions have their own pages and are not repeated here: the 10 percent brain myth and left brain versus right brain. The genetics of the same traits sits on the page on heritability. This page asks a narrower question: what do scans of people who also took cognitive tests show, how large are those effects, and which have survived being checked in larger samples.
2 What Did the 1988 PET Study Start, and What Is Neural Efficiency?
Modern neuroimaging of intelligence begins with a 1988 study of glucose use in the cortex, and its central idea, neural efficiency, says that more able people show lower brain activation on the same task. Haier and colleagues published Cortical glucose metabolic rate correlates of abstract reasoning and attention studied with positron emission tomography in the journal Intelligence in 1988. Haier and Jung's chapter describes it as the first of the neuroimaging studies of intelligence, the opening of what Haier calls Phase One of the field, a period that ran up to the 2007 P-FIT paper.
The hypothesis that grew from the study is stated plainly in the review by Neubauer and Fink, Intelligence and neural efficiency (Neuroscience and Biobehavioral Reviews, 2009). They describe the neural efficiency hypothesis as the claim that brighter individuals display lower, which is to say more efficient, brain activation while performing cognitive tasks, and they cite the 1988 study as its origin. They add that most early studies confirmed the hypothesis. The idea was memorable because it ran against the intuition that more able people must work their brains harder.
The phrase is easy to misread, so two clarifications belong here.
Lower activation is not a smaller brain or a lazy one. The hypothesis concerns how much of the brain's resources a given task recruits. It does not say that able people have less brain tissue, and the meta-analysis of brain volume described later finds the opposite association.
It is not the 10 percent myth in disguise. The popular saying that people use a tenth of their brain is false, as the dedicated page shows. Neural efficiency is about relative activation across people on a specific task, not about unused tissue.
Typical neuroimaging samples are small, a median of about 25 people according to Marek and colleagues, and the replication section below explains why small samples are a problem for any imaging claim. The right way to hold the 1988 finding is as the origin of a productive hypothesis whose later record is mixed, which the next section lays out.
3 Does Neural Efficiency Hold at Every Level of Difficulty?
It does not: Neubauer and Fink concluded that neural efficiency appears mainly when tasks are of low to moderate difficulty and mostly in frontal areas, and that on very complex tasks more able people tend to recruit more cortex, not less. The 2009 review is the main source used here because it brings together the studies that confirmed the hypothesis and the later ones that contradicted it. Their abstract reports that later research "revealed contradictory evidence or has identified some moderating variables," and it names four: sex, task type, task complexity and brain area.
Three findings from that review frame the current reading.
Task difficulty moderates the direction of the effect. Neubauer and Fink propose that neural efficiency may arise when people face tasks of subjectively low to moderate difficulty. In very complex tasks, they write, more able individuals seem to invest more cortical resources, which yields positive correlations between brain usage and cognitive ability. The sign of the association can therefore flip as a task gets harder.
Practice changes the picture. The review states that neuroscientific training studies suggest neural efficiency is also a function of the amount and quality of learning. Efficiency is seen on easier novel tasks, or after sufficient practice has allowed people to develop efficient strategies. So some of what looks like a fixed trait of the brain may reflect how well a task has been learned.
The effect is regional. Neubauer and Fink conclude that efficiency is mainly observable for frontal brain areas, not across the whole cortex.
Put together, the review says that a statement such as "smart people use less of their brain" is wrong as a general rule. A narrower statement survives: on easier, practiced or familiar tasks, higher ability often goes with lower activation in some frontal regions, and on very demanding tasks the relationship can reverse.
What does that imply for a test that mixes easy and hard items? The source papers do not say, and the following is our inference, not a finding: a test that ranges from easy to very hard items samples both regimes in one session, so activation recorded during a single task is a narrow window and not a portrait of intelligence. The reasoning tasks used to estimate fluid reasoning are described on the page on fluid intelligence tests, and the contrast with accumulated knowledge is on the page on fluid and crystallized intelligence.
4 What Is the P-FIT Model of Intelligence?
The Parieto-Frontal Integration Theory, P-FIT, is a 2007 synthesis of 37 neuroimaging studies that proposed intelligence draws on a distributed network linking parietal and frontal cortex, not on the frontal lobes alone. Rex Jung and Richard Haier published it in Behavioral and Brain Sciences under the title "The Parieto-Frontal Integration Theory (P-FIT) of intelligence: Converging neuroimaging evidence." Their review covered functional studies (functional MRI and PET) and structural ones (magnetic resonance spectroscopy, diffusion tensor imaging and voxel-based morphometry). They reported what they called a striking consensus: variations in a distributed network predict individual differences on intelligence and reasoning tasks.
The model names regions by Brodmann area, the numbered map of the cortex that neuroanatomists use. According to the abstract, it includes the dorsolateral prefrontal cortex (areas 6, 9, 10, 45, 46 and 47), the inferior parietal lobule (areas 39 and 40), the superior parietal lobule (area 7), the anterior cingulate (area 32), and regions in the temporal (areas 21 and 37) and occipital (areas 18 and 19) lobes. White matter regions, specifically the arcuate fasciculus, are also implicated. In their later chapter, Haier and Jung count 14 areas and stress the point the model was built to make: the regions relevant to intelligence are distributed throughout the brain, contrary to the view then popular that the frontal lobes alone are responsible.
The chapter also summarizes the information processing sequence the model suggests, which is worth reading as a proposal and not as an observation. The chapter attributes this five stage summary to Colom, Karama, Jung and Haier (2010).
Stage
Regions in the model
Proposed role
1
Occipital and temporal areas (extrastriate cortex, fusiform gyrus, Wernicke's area)
Process sensory information: recognition, imagery and elaboration of visual input, and analysis of the syntax of auditory input
2
Parietal areas 39, 40 and 7
Integrate and abstract the sensory information
3
Parietal areas interacting with frontal areas 6, 9, 10, 45, 46 and 47
Problem solving, evaluation and hypothesis testing
4
Anterior cingulate, area 32
Select a response and inhibit alternatives once the best solution is determined
5
White matter, especially the arcuate fasciculus
Reliable and efficient communication across the network
Two cautions apply. P-FIT was built from a qualitative review, a reading of what 37 studies tended to report, and not from a statistical pooling of effect sizes, which the chapter acknowledges when it contrasts it with later quantitative work. And the five stages are an inference from where the regions sit, not a measurement of a sequence in time. The chapter itself calls the model a source of testable hypotheses, for example that the sequence might repeat many times over seconds or minutes during reasoning and that its speed might differ among people.
Why did the model matter? Haier and Jung argue that it showed intelligence test scores are not meaningless artifacts, because quantifiable brain variables line up with them. That is a fair reading of what correlation can do: it supports construct validity, the claim that a score reflects something real about people. It does not show that any one region produces the score. The page on whether IQ is real treats the wider validity evidence, and the page on what IQ measures explains the construct the scans are being compared with.
5 Which Brain Regions Keep Appearing in Large Samples and Meta-Analyses?
Frontal regions recur in every design and parietal cortex appears in the meta-analytic and lesion evidence, but the largest sample also points to areas outside the P-FIT list, so that list is a starting map and not a complete one. Three lines of newer evidence can be compared, and each has a different design.
The first is quantitative meta-analysis. In their chapter, Haier and Jung describe the 2015 study by Ulrike Basten, Kirsten Hilger and Christian Fiebach, Where smart brains are different, in Intelligence. Where P-FIT was qualitative, the German group used a quantitative method across 28 studies with a combined sample of over 1,000 participants. According to the chapter, the findings supported a parieto-frontal network and added some other areas that might be involved, pending replication. Haier and Jung also report that roughly 100 further studies appeared after 2007, with larger samples and better methods, and that many of them supported P-FIT and its distributed nature. That second statement is the authors' own summary of a literature they helped create, so it is a reading by interested parties, even if a knowledgeable one.
The second is lesion mapping, which asks what is lost when a region is damaged. Gläscher and colleagues, in Distributed neural system for general intelligence revealed by lesion mapping (Proceedings of the National Academy of Sciences, 2010), studied 241 patients with focal brain damage. They derived g from a hierarchical factor analysis of several cognitive tasks and found significant associations between g and damage to a circumscribed but distributed network in frontal and parietal cortex, critically including white matter association tracts and frontopolar cortex. The authors suggest that general intelligence draws on connections between regions that integrate verbal, visuospatial, working memory and executive processes. Of the three lines of evidence, this one most directly supports the idea that some regions and fibers are necessary for performance, in this patient group.
The third is a very large structural sample. Cox and colleagues, in Structural brain imaging correlates of general intelligence in UK Biobank (Intelligence, 2019), analyzed 29,004 participants overall, aged 44 to 81, of whom 18,426 had both MRI and at least one cognitive test and at least 7,201 had a complete four test battery with MRI, depending on the brain measure. They derived a general factor from four varied cognitive tests, which accounted for one third of the variance in the test scores. Their abstract lists the largest regional correlates of g as volumes of the insula, frontal, anterior and superior and medial temporal, posterior and paracingulate, and lateral occipital cortices, plus thalamic volume and the white matter microstructure of thalamic and association fibers and of the forceps minor. Many of these regions made unique contributions and showed highly stable out of sample prediction.
The table sets the three lines of evidence side by side.
Study
Design
Sample
Regions or tracts reported
Basten and colleagues, 2015, as described by Haier and Jung
Quantitative meta-analysis of functional and structural studies
28 studies, over 1,000 participants combined
A parieto-frontal network, plus additional areas that need replication
Gläscher and colleagues, 2010
Voxel-based lesion-symptom mapping
241 patients with focal brain damage
A distributed frontal and parietal network, including white matter association tracts and frontopolar cortex
Cox and colleagues, 2019
Cross-sectional structural MRI, latent g from four tests
29,004 participants overall, ages 44 to 81
Insula, frontal, temporal and occipital cortices, thalamus, and thalamic and association fiber microstructure
Comparing the Cox list with the Brodmann areas named in P-FIT is our own comparison of two published lists, and it shows partial overlap. Frontal and temporal cortex appear in both. The insula and the thalamus, which appear among the Cox regional correlates, are not in the P-FIT list as given in the abstract. That does not refute P-FIT, because the two studies used different measures and asked a different statistical question, but it does support the note from the meta-analysis that the full list is still being established.
What can be said with confidence is modest: frontal cortex and the white matter linking regions appear across designs, parietal cortex appears in the meta-analytic and lesion work, and no single region is the seat of intelligence.
6 How Strongly Is Brain Volume Related to IQ?
Total brain volume is positively associated with IQ, but the best estimates put the correlation near r = .2 to .3, which leaves most of the variance unexplained. The topic has its own treatment elsewhere on the site, in the section on biological correlates in the page on whether IQ is real, so this section gives only the figures needed to compare brain size with the network findings.
The central source is the 2015 meta-analysis by Pietschnig, Penke, Wicherts, Zeiler and Voracek, Meta-analysis of associations between human brain volume and intelligence differences (Neuroscience and Biobehavioral Reviews). They identified 88 studies examining 148 healthy and clinical mixed sex samples with more than 8,000 individuals. The pooled association was r = .24, which corresponds to R squared of .06, about 6 percent of the variance. The association generalized over age (children versus adults), over IQ domain (full scale, performance and verbal IQ) and over sex. The authors also applied methods for detecting publication bias, and they report that strong positive coefficients have been reported frequently while small and nonsignificant ones appear to have often been omitted. In their words, the strength of the association has been overestimated in the literature but remains robust after accounting for several types of dissemination bias, and reported effects have been declining over time. They add that it is not warranted to treat brain size as an isomorphic proxy of intelligence differences.
The second source is a study designed to avoid those problems. Nave, Jung, Karlsson Linnér, Kable and Koellinger, in Are Bigger Brains Smarter? Evidence From a Large-Scale Preregistered Study (Psychological Science, 2019), preregistered their analysis in a new sample of 13,608 adults from the United Kingdom. The sample was about 70 percent larger than the combined samples of all earlier investigations, and the analyses controlled for sex, age, height, socioeconomic status and population structure. They found a robust association between total brain volume and fluid intelligence of r = .19, and one between total brain volume and educational attainment of r = .12. The associations were mainly driven by gray matter rather than white matter or fluid volume, and they were similar for both sexes and across age groups.
The third figure comes from the UK Biobank paper already cited. Cox and colleagues report r = 0.276, with a 95 percent confidence interval of 0.252 to 0.300, between age and sex corrected total brain volume and a latent general factor. They also report that a model combining several global measures of gray and white matter structure accounted for more than double the g variance in older participants compared with middle aged ones, 13.6 percent against 5.4 percent.
Source
Design
Sample
Measures
Result
Pietschnig and colleagues, 2015
Meta-analysis with publication bias checks
88 studies, 148 samples, over 8,000 people
Brain volume and IQ
r = .24, R squared = .06
Nave and colleagues, 2019
Preregistered, controlled for sex, age, height, socioeconomic status and population structure
13,608 UK adults
Total brain volume and fluid intelligence
r = .19 (educational attainment r = .12)
Cox and colleagues, 2019
Cross-sectional UK Biobank analysis
29,004 participants overall, ages 44 to 81
Total brain volume and latent g from four tests
r = 0.276, 95 percent interval 0.252 to 0.300
Squaring a correlation gives the share of variance it accounts for, and that arithmetic is ours, not the authors'. For r = .19 it is about 3.6 percent, and for r = .276 about 7.6 percent. Two people of identical brain volume can differ widely in IQ, and two people with identical IQ can differ widely in volume. The pattern is a tendency across a population. It cannot be applied to an individual without enormous error, and it says nothing about how the relationship arises. Possible routes include shared genes, nutrition and health in development, education and age related change, and the observational designs here cannot separate them.
Two nearby pages show why the caution matters. The page on IQ and height discusses how body size and growth timing complicate any link with brain size. The page on alcohol and IQ explains why a smaller brain volume among heavy drinkers is not the same thing as a measured loss of ability, since the direction of cause is open.
7 What Does White Matter Add to the Picture?
White matter integrity goes with intelligence differences, and the study reviewed here ties the association to processing speed. Gray matter holds the regions. White matter holds the fibers that carry signals between them, and the P-FIT model names the arcuate fasciculus for that reason.
Penke and colleagues studied this in Brain white matter tract integrity as a neural foundation for general intelligence (Molecular Psychiatry, 2012). They scanned 420 older adults in their early 70s and measured three biomarkers of integrity, fractional anisotropy, longitudinal relaxation time (T1) and magnetization transfer ratio, in 12 major tracts. The tracts were correlated enough to extract three general factors of brain-wide integrity, one for each biomarker. Each was independently associated with general intelligence, and together they explained 10 percent of the variance. Their effect was completely mediated by information processing speed. The authors interpret this as a functionally plausible model: structurally intact fibers across the brain provide the infrastructure for fast information processing within widespread networks, which supports general intelligence.
Reading that result carefully matters, because "mediated" is a statistical term. In a single cross-sectional sample of older adults it means the data fit a model in which tract integrity relates to speed and speed relates to intelligence, and no experiment changed anyone's white matter. The finding is therefore an association with a proposed mechanism. The page on processing speed covers why speed is one of the first abilities to change across adulthood, and the same ability is one of the six ACIS indices, but the Penke study does not tell you that a slower score on any test reflects a specific tract.
White matter appears in the other sources too. Gläscher and colleagues' lesion map critically included white matter association tracts. Cox and colleagues list the white matter microstructure of thalamic and association fibers and of the forceps minor among the largest regional correlates of g. Haier and Jung's chapter adds that integrity of white matter fibers throughout the brain is related to IQ in studies using newer imaging methods, and it describes twin work reporting that white matter integrity is most heritable in frontal and parietal areas, with IQ correlated with some of those tracts. That heritability claim is the chapter's summary of twin studies it cites, and this page has not checked it against the original twin papers. The genetic side is the subject of the page on heritability.
8 What Do Network and Connectivity Studies Show?
Network studies treat intelligence as a property of how brain regions communicate, and the strongest result is that resting-state connectivity predicts intelligence scores in new people, though no single connection or hub carries the prediction. The shift from regions to networks follows from the white matter results above: if fibers matter, then the pattern of connections should matter too.
Haier and Jung's chapter explains the method behind much of this work, graph analysis. Each brain area becomes a node, the connections between areas become links, and the number and strength of links can be quantified. Areas with many connections to others are called hubs. The chapter notes that, until recently, attempts to predict IQ from brain images with multivariate statistics were not cross-validated in independent samples, and that not all brains seem to work the same way, so a set of areas that matters for one person may not matter for another. It describes an early graph analysis of 207 individuals in which IQ scores were related to connections among many areas, including regions identified in P-FIT.
Three developments define the current picture.
A theory. Aron Barbey's Network Neuroscience Theory of Human Intelligence (Trends in Cognitive Sciences, 2018) is an opinion article that surveys the evidence and proposes that g emerges from the small-world topology of brain networks and from the dynamic reorganization of their community structure in the service of system-wide flexibility and adaptation. It is a theoretical synthesis. The article does not report a new experiment, so it is a framework to be tested, not a result.
Connectivity fingerprints. The chapter summarizes a Human Connectome Project study by Finn and colleagues, published in 2015, in which 126 people were scanned in six sessions that included resting and task conditions. Connectivity patterns among 268 brain nodes in 10 networks were stable within a person across sessions and unique enough to identify the person, and these individual patterns predicted fluid intelligence, with the strongest correlations in frontoparietal networks. Haier and Jung call it, in their view, the first study with an independent sample for cross-validation to predict IQ from functional MRI. The summary is theirs, and the 2015 paper itself is not among the sources opened for this page.
Large sample prediction. Dubois, Galdi, Paul and Adolphs, in A distributed brain network predicts general intelligence from resting-state human neuroimaging data (Philosophical Transactions of the Royal Society B, 2018), used the final release of the Young Adult Human Connectome Project, with 884 subjects after exclusions and a full hour of resting-state fMRI each. They controlled for gender, age and brain volume, derived a general intelligence estimate from multiple cognitive tasks and, with a cross-validated predictive framework, predicted 20 percent of the variance in general intelligence from resting-state connectivity matrices. No single anatomical structure or network was responsible or necessary for the prediction, which relied on redundant information distributed across the brain. They also state that the most replicated neural correlate of intelligence to date is total brain volume, a coarse measure that says little about function.
The 2022 systematic review by Vieira and colleagues, On the prediction of human intelligence from neuroimaging (Intelligence), put numbers on the whole prediction literature. They assessed 37 studies that used machine learning with structural, functional or diffusion MRI to predict intelligence in cognitively normal people, a different set from the 37 behind P-FIT. In a meta-analysis, studies predicting general intelligence from Human Connectome Project functional MRI averaged r = 0.42, with a 95 percent confidence interval of 0.35 to 0.50, while studies predicting fluid intelligence averaged r = 0.15, with an interval of 0.13 to 0.17. The authors conclude that the quality of the measurement of intelligence moderates the association, and they report that failure to treat confounders and small sample sizes were common problems that raised the risk of bias. They describe the literature as reaching maturity and note that extending findings to new populations is imperative.
Two points follow. First, the same data give very different answers depending on the cognitive measure: a score built from many tasks was predicted far better than a fluid reasoning score, which echoes the opening point that g is estimated and that the estimate limits what any brain measure can explain. Second, predicting 20 percent of the variance in a held-out group is real and also modest. On one scale, and this arithmetic is ours, a correlation of 0.42 corresponds to about 18 percent of variance and 20 percent of variance to a correlation of about .45. Both leave most variation unexplained.
9 Why Do So Many Brain and Intelligence Findings Fail to Replicate?
Many early brain and intelligence findings did not replicate largely because they came from samples of a few dozen people, and in brain-wide association studies the true effects are much smaller than those samples can detect without inflation. The key demonstration is Marek and colleagues, Reproducible brain-wide association studies require thousands of individuals, published in Nature in 2022.
Marek and colleagues define a brain-wide association study, BWAS, as a study of the association between common individual variability in brain structure or function and cognition or psychiatric symptoms. They note that the median neuroimaging study sample size is about 25. To test whether that is enough, they used three of the largest neuroimaging datasets available, the Adolescent Brain Cognitive Development study (11,874 participants in the paper's description), the Human Connectome Project (1,200) and UK Biobank (35,735), a total of around 50,000 individuals. They drew subsamples from 25 up to 32,572 people. The key results from their main sample of 3,928 people are below.
The typical effect is tiny. Across all brain-wide associations, which paired cortical thickness or resting-state connectivity with 41 measures of demographics, cognition and mental health, the median univariate effect size was an absolute r of 0.01. The top 1 percent of around 11 million possible associations exceeded an absolute r of 0.06. The largest correlation that replicated out of sample was 0.16.
Small studies report bigger numbers. Smaller brain-wide association studies have reported larger univariate correlations, r greater than 0.2, than the largest effects measured in the much larger samples. With 25 participants, the 99 percent confidence interval for these associations was r plus or minus 0.52, and two independent subsamples of that size could reach opposite conclusions about the same association from sampling variability alone.
Inflation persists in the thousands. In split halves of 1,964 people each, the top 1 percent largest effects were still inflated by r = 0.07, which is 78 percent, on average.
Power needs thousands. For the top 1 percent largest effects (r greater than 0.06), the paper reports that 9,500 participants were required for 80 percent power under a strict multiple comparison threshold, and 2,200 under an uncorrected threshold of .05.
Publication selects the inflated results. At small sample sizes, the largest and most inflated effects are the most likely to be statistically significant and, the authors write, paradoxically the most likely to be published.
The abstract also says which effects held up better: functional MRI more than structural MRI, cognitive tests more than mental health questionnaires, and multivariate methods more than univariate ones. And it contrasts brain-wide association studies with approaches that have larger effects, specifically lesions, interventions and within person designs. Multivariate methods are the ones that give most of the prediction results reported above, which fits the abstract's finding that multivariate effects were more robust than univariate ones.
What does this mean for the findings above? A few caveats keep the reading honest. The median of 0.01 describes brain-wide associations across many phenotypes, including mental health, so it is not a measure of how strongly the brain relates to intelligence in particular. The brain volume figures from Nave and Cox, r = .19 and r = 0.276 in samples of 13,608 and 29,004, were obtained in samples in the thousands, the size range the paper says such studies need, and they are one global measure per person, not millions of local tests, which is our observation and not the authors'. Lesion studies, where damage is large, are the other kind of design the abstract describes as having larger effects. The findings most at risk are the specific ones: a particular region, a particular connection or a particular hub reported in a study of a few dozen people. Studies of that kind, with samples near the median of about 25, are the ones most exposed to the problem.
10 What Survives the Replication Test, and What Does Not?
Three findings survive in large samples, namely a modest brain volume association, a distributed parieto-frontal pattern and above chance prediction from connectivity, while region level and hub level claims from small studies remain unconfirmed. The table applies one rule to each finding: how large was the best sample, and what effect size did the source print? The strength labels are our reading of the sources on this page and not a verdict from any of the authors.
Finding
Best evidence on this page
Effect size printed
Our strength label
Larger total brain volume goes with higher IQ
Meta-analysis of 88 studies (Pietschnig); preregistered sample of 13,608 (Nave); 29,004 participants (Cox)
r = .24, r = .19 and r = 0.276
Robust, and modest
Frontal and parietal association cortex are involved, as part of a distributed network
37 study review (Jung and Haier); meta-analysis described in the chapter (Basten); 241 lesion patients (Gläscher); frontal and other regions in UK Biobank (Cox)
No pooled regional effect size is printed in the sources opened
Robust as a pattern, unquantified region by region
White matter integrity goes with intelligence, through speed
420 older adults (Penke), with support from lesion and UK Biobank findings
10 percent of variance, fully mediated by processing speed in the model
Supported, but the headline figure rests on one sample
Lower activation in more able people (neural efficiency)
Review of the activation literature (Neubauer and Fink)
No pooled effect size in the abstract
Mixed: depends on difficulty, practice, sex and region
Resting-state connectivity predicts intelligence in new people
884 people (Dubois); systematic review of 37 studies (Vieira)
20 percent of variance; r = 0.42 for general intelligence and r = 0.15 for fluid intelligence
Supported, and sensitive to how intelligence is measured
A particular region, edge or hub is the key to intelligence
Brain-wide association studies in about 50,000 people (Marek)
Median absolute r of 0.01, top 1 percent above 0.06, largest replicated 0.16
Not replicated at typical sample sizes
Small-world topology and flexible communities give rise to g
Theory article (Barbey)
None: an opinion article
A hypothesis to be tested
One observation from the table is worth stating plainly: robustness and size are separate questions. Brain volume is the most replicated correlate, and its effect is still only about 4 to 8 percent of variance by our arithmetic. A finding can be real and small.
No: scans predict intelligence scores in groups with modest accuracy, and the IQ of an individual is still measured by a cognitive test, not read from an image. Haier has argued, in the passage the chapter quotes, that a brain image would not be sensitive to factors such as motivation or anxiety that can influence test scores. The chapter raises the idea as an empirical question, and it also states the problem: the goal of predicting IQ from neuroimages has been elusive, because individual differences in both intelligence and brain characteristics are wide, and attempts that claimed success mostly failed to cross-validate in independent samples until recently.
The best current numbers are the ones already reported. In the Dubois study, resting-state connectivity predicted 20 percent of the variance in general intelligence in held-out people. In the Vieira review, studies using Human Connectome Project functional MRI averaged r = 0.42 for general intelligence and r = 0.15 for fluid intelligence. A simple calculation shows what those correlations mean for one person, and it is our arithmetic, not a figure from either paper. Assume a straight line relationship, normally distributed scatter and the usual IQ standard deviation of 15, which the page on standard deviation 15 explains. Then a correlation of 0.42 leaves a typical prediction error of about 14 IQ points, and a correlation of 0.15 leaves about 15, essentially the spread of IQ in the population. Under the same assumptions, a 95 percent range around a prediction made with a correlation of 0.42 would span roughly 27 points either way. That is too wide to classify anyone, and it applies even to the best studies reviewed.
Several other limits keep scans from replacing tests.
Age range. Penke's sample was in its early 70s, and the UK Biobank sample of Cox and colleagues was 44 to 81. Cox's multivariable model explained 13.6 percent of the variance in older participants and 5.4 percent in middle age, so relationships may differ at other ages. The page on how IQ changes with age covers the score side.
Measurement. Every brain prediction is trained on a cognitive test score. Vieira's meta-analysis shows that prediction quality depends on the quality of that score, so the scan depends on the test and cannot replace it.
Design. Most results are cross-sectional associations, so they cannot say why a pattern exists or whether it would change if the brain changed.
The causal vocabulary on this page follows the designs. Volume and connectivity studies show that measures go together, and prediction studies show that a model trained on some people forecasts scores in others. Lesion mapping shows that damage in certain regions goes with lower scores, which is the closest the evidence comes to necessity. The white matter mediation model is a statistical account of one sample. None of the studies randomly changed a brain feature and then measured intelligence. Whether intelligence scores can be raised by training or schooling is a different literature, summarized on the page about whether you can increase your IQ. Whether brain features pass from mother or father belongs to the page on that question, and what intelligence is in the first place is covered in What Is Intelligence?.
12 What Do Brain Findings Mean for Reading Your Own Score?
Brain research supports reading an IQ score as an estimate of something real, and the Standards for Educational and Psychological Testing say that any such estimate needs a stated interpretation and validity evidence, with due weight given to findings that cut against it. The Standards for Educational and Psychological Testing, published jointly by AERA, APA and NCME in 2014, open their chapter on validity with Standard 1.0: clear articulation of each intended test score interpretation for a specified use, with appropriate validity evidence in support of each. The comment to Standard 1.2 asks that a presentation of evidence give due weight to all relevant findings in the scientific literature, including those inconsistent with the intended interpretation. That is the same discipline the replication results above demand of brain studies, and it applies to a test publisher just as much as to a laboratory.
Three further passages bear on how to read a score next to this literature.
Relations to other variables. Standard 1.16 says that when validity evidence includes data on other variables, the rationale for choosing them should be given, along with evidence on their technical quality. A brain measure could in principle be one such variable, but only with that rationale and quality evidence. Marek and colleagues show why a small imaging sample would be a weak basis for such evidence.
The form of an association. The comment to Standard 1.18 says evidence about the overall association between variables should be supplemented by information about the form of that association and about its variability in different ranges of test scores. A correlation of .24 is an overall association, and it hides how much individuals scatter around it.
Context. The APA Guidelines for Psychological Assessment and Evaluation, approved in March 2020, state in the rationale to Guideline 7 that individual performance on psychological tests is only one piece of an assessment, conceptualized within a context and secured from multiple sources. The rationale to Guideline 5 adds that validity is not an inherent property of a test but refers to the degree to which evidence and theory support its use for a particular purpose.
What does this add up to for someone who takes an IQ test? Four practical points follow, and they are our synthesis of the passages and of the brain evidence, not a list printed in either document.
Treat the score as a position, not a measurement of a brain. A Full Scale IQ places a person relative to a reference group. The page on how IQ scores are normed explains that logic and the page on score versus percentile explains why the two scales differ.
Use the interval. Brain studies show how noisy even good measures are, and a test score is noisy too. Look for the confidence interval, not only the number. The page on reliability and validity explains the difference between a score that is repeatable and a score that means what it is said to mean.
Do not borrow brain findings as proof of a product. An online test cannot claim brain-based validity because brain research exists. A product earns trust through its own documentation, which for ACIS is in the technical manual.
We sell a paid online assessment, so read this as a disclosure. ACIS has 20 subtests in six domains of the Cattell-Horn-Carroll model, described on the page about the CHC model, and reports a Full Scale IQ and six primary indices on the standard scale, with a mean of 100 and a standard deviation of 15, and the report gives percentiles and a 95 percent confidence interval. Scaled subtest scores run from 1 to 19. Adult norms cover ages 16 to 90. The assessment is online and unsupervised, not a clinical or diagnostic instrument, not suitable for hiring, school accommodations or admission to high IQ societies, and available in English only. Nothing on this page implies that it reads the brain, and the page on choosing a test and reading the result and the page on how to tell whether a score is real set out what to look for in any test.
Every figure on this page comes from one of the sources below. DOIs were checked against Crossref on October 6, 2026. The abstracts were read through PubMed and Europe PMC on October 6, 2026, and the Marek full text through Europe PMC; the Standards and the APA Guidelines were downloaded and searched on the same date. The 1988 PET study and the Basten meta-analysis were not available to us in full, and their descriptions here come from the Neubauer and Fink review and from Haier and Jung's chapter, as stated in the text.
Haier, R. J., and Jung, R. E. The Parieto-Frontal Integration Theory: Assessing Intelligence from Brain Images, chapter 7. In D. P. Flanagan and E. M. McDonough (Eds.), Contemporary Intellectual Assessment: Theories, Tests, and Issues, 4th ed. Guilford Press, 2018.
American Educational Research Association, American Psychological Association and National Council on Measurement in Education. Standards for Educational and Psychological Testing. AERA, 2014, chapter 1, Standards 1.0, 1.2, 1.16 and 1.18, read October 6, 2026.
American Psychological Association, APA Task Force on Psychological Assessment and Evaluation Guidelines. APA Guidelines for Psychological Assessment and Evaluation. Approved by the APA Council of Representatives in March 2020, Guidelines 5 and 7, read October 6, 2026.
14 Frequently Asked Questions
What part of the brain controls intelligence?
No single part does. Imaging studies link intelligence scores to a distributed network, with frontal and parietal cortex most often reported, plus the white matter fibers connecting them. Even in very large samples, each region or measure explains only a small share of score differences, so intelligence is better described as network wide than as located.
Where is intelligence located in the brain?
Intelligence is not stored in one location. The P-FIT model of 2007 describes a network spanning frontal, parietal, temporal and occipital regions, and lesion mapping in 241 patients found a distributed frontal and parietal system. Researchers describe where scores correlate with brain features, not a place where intelligence lives.
Do smarter people use less of their brain?
Not as a general rule. The neural efficiency hypothesis says more able people show lower activation on the same task, and a 2009 review found this mainly on easier tasks and in frontal areas. On very complex tasks, more able people tended to recruit more cortex. The popular claim that anyone uses only 10 percent is a myth.
What is the P-FIT model of intelligence?
P-FIT is the Parieto-Frontal Integration Theory, published by Jung and Haier in 2007 after reviewing 37 neuroimaging studies. It proposes that intelligence draws on a distributed network of parietal, frontal, temporal and occipital regions and the white matter linking them, with sensory processing, integration, problem solving and response selection as proposed stages.
Is intelligence related to brain size?
Yes, modestly. A meta-analysis of 88 studies found r = .24 between brain volume and IQ, about 6 percent of variance, and a preregistered sample of 13,608 UK adults found r = .19 with fluid intelligence. People of the same brain volume still differ widely in IQ, and the data cannot show cause.
Which brain networks are linked to intelligence?
Frontoparietal networks are the most often named. A Human Connectome Project study summarized by Haier and Jung found the strongest links to fluid intelligence there, and an 884 person study predicted 20 percent of the variance in general intelligence from resting connectivity spread across the brain, with no single network necessary.
Can a brain scan measure IQ?
Not for an individual. Models trained on connectivity predict scores in groups, averaging r = 0.42 for general intelligence in one systematic review, which leaves most variation unexplained. By our arithmetic that implies typical errors near 14 IQ points. A cognitive test remains the way IQ is measured.
What is the neural efficiency hypothesis?
It is the idea that brighter individuals show lower, more efficient brain activation while performing cognitive tasks. The 1988 PET study by Haier and colleagues is its cited origin. A 2009 review found support mainly for easier tasks, after practice and in frontal areas, with contradictory evidence elsewhere and moderators such as sex.
How big is the correlation between brain volume and IQ?
About r = .2 to .3. Pietschnig and colleagues reported r = .24 across 88 studies, Nave and colleagues r = .19 in a preregistered sample, and a UK Biobank analysis r = 0.276. Squared, these are roughly 4 to 8 percent of variance, and the 2015 meta-analysis cautions against treating brain size as a proxy for intelligence.
What did Marek and colleagues find in 2022?
Using about 50,000 people, they found brain-wide associations with cognition and mental health were far smaller than assumed: a median absolute r of 0.01 and a largest replicated r of 0.16 in the main sample. At typical sample sizes studies were underpowered and inflated, and reproducibility required thousands of participants.
What does white matter have to do with intelligence?
White matter fibers connect brain regions, and better integrity goes with higher scores. In 420 older adults, three white matter factors together explained 10 percent of the variance in general intelligence, and the model had processing speed carry the whole effect. Lesion and UK Biobank studies also implicate association tracts.
What did the lesion study of 241 patients show?
Gläscher and colleagues found that damage to a circumscribed but distributed frontal and parietal network, including white matter association tracts and frontopolar cortex, was associated with lower general intelligence. The authors suggest g draws on connections between regions that integrate verbal, visuospatial, working memory and executive processes. It studied patients with focal damage.
Why do brain imaging studies of intelligence fail to replicate?
Mostly because samples were small and true effects are modest. With a median neuroimaging sample of about 25, chance can produce correlations above 0.2 that disappear in larger samples, and significant results are the ones published. Larger samples, preregistration and independent replication reduce inflation but do not make small effects large.
What is resting-state connectivity and can it predict intelligence?
Resting-state connectivity is the statistical coupling of activity between brain regions while a person lies still in a scanner. In 884 people it predicted about 20 percent of the variance in general intelligence. Predictions of fluid intelligence alone were far weaker, averaging r = 0.15 in the systematic review, so measurement quality matters.
Does a bigger brain mean a higher IQ for me?
Not reliably. The association is a tendency across populations, with correlations near .2 to .3, so brain volume tells almost nothing about one person's score. People with the same volume can have very different IQs. No scan result can replace testing or be used to rank you.
Does brain research show that intelligence can be changed?
Not directly. The studies on this page are mostly observational, relating brain measures to scores at one time, so they cannot show that changing a brain feature changes intelligence. Whether training or schooling changes scores is a separate question, answered by intervention studies rather than by brain scans.
Is intelligence in the frontal lobes?
Not only. The frontal lobes feature in the P-FIT model and in lesion and large sample studies, but Jung and Haier built the model to show that relevant regions are distributed, contrary to the then popular view that the frontal lobes alone were responsible. Parietal cortex and white matter also appear repeatedly.
Do IQ scores reflect real brain differences?
At the group level, yes. Brain volume, white matter, regional structure and connectivity all correlate with intelligence scores, which Haier and Jung argue shows the scores are not meaningless artifacts. The correlations are modest, so a score is not a readout of any single brain feature, and it carries measurement error.
How should I read an IQ score in light of brain research?
As an estimate with a reference group and an interval. Brain studies support the idea that scores reflect something real, but their effects are small and mostly group level. The Standards for Educational and Psychological Testing ask for validity evidence for each interpretation, so read a score with its percentile and confidence interval.
Can brain scans diagnose low or high intelligence?
No. The research on this page describes groups, and the sources report modest correlations and prediction errors far too large for classifying individuals. Nothing here is a diagnostic claim. Questions about learning, development or cognitive concerns belong with a qualified clinician using standardized assessment.
Does ACIS use brain imaging?
No. ACIS is an online, self administered cognitive assessment with 20 subtests that reports a Full Scale IQ and six index scores, with percentiles and a 95 percent confidence interval. It measures performance on tasks, it is not a clinical instrument, and it does not scan or read the brain.
Take the assessment
You get a profile, not a number
ACIS measures six CHC domains across 20 subtests and reports each one with its own normed score and confidence interval, so you can see where you are strong and where you are not.