Average IQ by Language Why the Question Breaks, and What Language Really Does
No language makes anybody more or less intelligent. But the language you are tested in changes your measured score substantially, and on one subtest the effect comes down to how many syllables your number words happen to have.
0 Quick Answer
Updated August 16, 2026 by Structural. There is no meaningful average IQ by language. Languages are not populations, speakers of any major language span the full ability range, and no research programme has ever framed the question this way because it does not correspond to anything measurable.
Direct answer: what does exist, and is well documented, is that language affects measurement. Being tested in a second language depresses verbal scores substantially. Test translation is not translation but re-standardisation. And on at least one subtest, the structural properties of a language change scores directly, independent of anybody's ability.
That last finding is the most striking in this area. Digit span performance depends partly on how long it takes to say the number words, so speakers of languages with short number words recall longer sequences. It is a property of the vocabulary, not of the person, and it has been demonstrated within bilinguals tested in both of their own languages.
Several things go wrong at once when this question is asked, and separating them shows why no dataset could answer it.
A language is not a population. English is spoken by well over a billion people across every continent, at every level of education, in every economic circumstance. Whatever the average of that group is, it is not a fact about English, and the group has no coherence beyond the shared language.
Language and country do not align. Most large languages are spoken across many countries with wildly different educational systems and economic conditions, and most countries contain speakers of many languages. Any language-level average would be a weighted mixture of national conditions, dominated by whichever populations happened to be largest.
The measurement is circular for verbal ability. A verbal test measures ability as it has developed within a specific language, so comparing verbal scores across languages compares different instruments. There is no language-neutral verbal measure, and constructing one is not a technical problem awaiting solution but a contradiction in terms.
And nobody collects it. There is no research programme measuring cognitive ability by language, because the question does not correspond to a construct anybody has proposed. What exists instead is a substantial literature on how language affects measurement, which is a different and answerable question.
This is the same structure as the state and sport versions of the question, covered in Average IQ by State: a group label that does not define a sampleable population, attached to a measure that was not designed to compare across it.
2 The Number Word Effect
The clearest demonstration that language affects scores independently of ability comes from digit span, and it is worth explaining in full because it is unusually decisive.
Digit span asks you to hear a sequence of digits and repeat it back. Performance depends on a rehearsal mechanism that maintains material by repeating it, either aloud or internally, and that mechanism is time-limited. What determines how much you can hold is roughly how much you can say in the available window.
It follows that if number words are shorter to pronounce, more of them fit in the same window, and span is longer. This is not speculation, it is a direct prediction from the model, and it has been tested.
The decisive study compared Welsh and English speakers. Welsh number words take longer to pronounce than their English equivalents, and Welsh-English bilinguals showed shorter digit spans when tested in Welsh than in English. The same people, the same working memory, a different measured span, driven entirely by how long the words take to say.
The comparison within bilinguals is what makes this conclusive. Comparing Welsh speakers with English speakers would confound language with everything else that differs between the groups. Comparing the same individuals in both of their own languages removes every such confound and leaves only the linguistic property.
The same mechanism explains a widely noticed pattern. Chinese number words are unusually short, close to monosyllabic, and digit spans measured in Chinese are correspondingly longer than in most European languages. This is regularly cited as evidence of a group difference in working memory. It is evidence about phonology, and the within-bilingual design shows exactly why.
The general lesson is that a subtest can be sensitive to a structural property of the language it is administered in, and that norms are therefore language-specific in ways that are not obvious from the task description. Digit span is described in The Subtests Inside an IQ Test.
3 Testing in a Second Language
This is the largest practical effect in this area and the one that affects the most people.
Somebody assessed in a language they did not grow up speaking scores below their ability on verbal measures, and the size of the gap depends on when they acquired the language, how much they use it, and in which contexts. The effect is largest on vocabulary and knowledge measures, which depend on cumulative exposure, and it is substantial enough to change how a profile reads.
The reason is that verbal subtests measure ability as developed through a specific linguistic environment. Vocabulary depth accumulates over decades of reading, conversation, and formal instruction. Somebody with twenty years of exposure to a language has a smaller vocabulary in it than a native speaker of the same age, regardless of their reasoning ability, and no statistical adjustment recovers the difference.
The distortion is not uniform across the battery, which is what makes it dangerous rather than merely inconvenient. Verbal indices fall the most. Working memory falls somewhat, because verbally presented material must be processed in the weaker language. Fluid reasoning and visual spatial measures fall least, since they depend on instructions rather than content. Processing speed is affected mainly through instruction comprehension.
The resulting profile shows a large gap between verbal and nonverbal indices, which is a pattern that in a monolingual examinee can indicate a specific learning difficulty or a language disorder. In a second-language examinee it indicates that they were tested in their second language. Distinguishing the two requires knowing the language history, which is one of several reasons that intake interviews exist and that a report without background information is not an assessment.
This is also the confound that distorts the field rankings in Average IQ by College Major, where disciplines enrolling many international students see their verbal means depressed for reasons that have nothing to do with the discipline.
4 Why Translating a Test Is Not Translating
Adapting a cognitive battery to another language is a research project, not a translation job, and the reasons illustrate what a test actually is.
Item difficulty does not survive translation. A vocabulary item calibrated as moderately difficult in one language may be trivially easy or impossibly obscure in another, because word frequency and acquisition age differ. The careful difficulty ordering that lets a subtest discriminate across the ability range is destroyed by direct translation.
Some items have no equivalent at all. Items relying on idiom, wordplay, phonology, or culturally specific knowledge cannot be translated in any useful sense and must be replaced with new items serving the same measurement function, which then require their own calibration.
Norms do not transfer. Even a perfectly adapted instrument must be standardised on a representative sample of the new population, because a standard score is meaningless against a reference group that did not take this version of the test. This is the expensive part, and it is why adaptations lag original publications by years.
Structural equivalence must be demonstrated rather than assumed. The question of whether the adapted instrument measures the same constructs in the same way is empirical, tested through analyses of measurement invariance, and it sometimes fails. When it does, scores from the two versions are not comparable no matter how careful the translation was.
The International Test Commission publishes detailed guidelines covering this process precisely because doing it badly is common and the consequences are invisible in the resulting scores. A poorly adapted instrument produces numbers that look exactly like good ones.
The practical implication for anybody assessed in a language other than the instrument's original is to ask which adaptation was used and when it was standardised locally. That date, rather than the edition number, determines how current the norms behind the score are, for the reasons in Average IQ by Generation.
5 What Bilingualism Actually Does
The effects of bilingualism on cognition have been researched extensively and contested vigorously, and the honest summary distinguishes what is settled from what is not.
Reasonably settled: vocabulary in each language is smaller. Bilinguals typically have a smaller vocabulary in each of their languages than monolinguals have in their single one, while their total vocabulary across both is comparable or larger. This matters directly for testing, because a vocabulary subtest samples one language and therefore understates a bilingual's lexical knowledge as a whole.
Reasonably settled: retrieval is slightly slower. Bilinguals show small delays in naming and lexical access, consistent with managing two active lexical systems. The effect is small and appears on timed verbal tasks.
Heavily contested: the executive function advantage. Influential earlier work proposed that constantly managing two languages trains inhibitory control and cognitive flexibility, producing a general executive advantage. Subsequent attempts to replicate produced mixed results, and systematic reviews have questioned whether the effect exists at all, pointing to small samples, publication bias favouring positive findings, and inconsistent task selection.
The current state is that the advantage, if present, is smaller and more conditional than originally proposed, and the field has not converged. Anybody encountering a confident claim in either direction is reading a partisan account of an unresolved question.
What is not in dispute is the measurement consequence. A bilingual tested in one language produces a verbal score that understates their linguistic knowledge, and that is true regardless of how the executive function debate resolves. The measurement issue is practical and immediate, and the executive question is theoretical and open.
6 Do Nonverbal Tests Solve It?
The obvious response to all of this is to use tests without verbal content. That helps substantially and does not solve the problem, and understanding why is useful.
Nonverbal and language-reduced instruments exist precisely for this situation, using matrix reasoning, pattern completion, and spatial tasks with minimal verbal instruction. They are the correct choice when language would otherwise confound the assessment, and they are widely used for exactly that reason.
Three limitations remain. Instructions still require comprehension, and a misunderstood instruction costs items regardless of how nonverbal the content is. Demonstration and gesture reduce this but do not eliminate it.
Familiarity with the test format is not culturally neutral. Matrix reasoning assumes comfort with abstract figural puzzles presented on paper or screen, and that comfort is a product of schooling of a particular kind, as covered in Average IQ by Generation. Somebody who has never encountered this genre of task is disadvantaged in a way that has nothing to do with reasoning ability.
And the coverage is narrower. A nonverbal instrument measures fluid reasoning and visual processing well and does not measure crystallised knowledge at all. That is a substantial part of cognitive ability simply absent from the assessment, so a nonverbal score is not equivalent to a full scale score and should not be reported as one.
There is a further subtlety about what a nonverbal score can be compared to. If somebody is assessed nonverbally because language would confound a full battery, their score cannot be set against a full scale figure from anybody else, because the two summarise different sets of abilities. Reports handle this by presenting the nonverbal composite under its own label with an explicit statement of what it excludes, and a report that quietly presents it as a general ability score has misrepresented what was administered.
The honest position is that nonverbal testing removes the largest confound at the cost of construct coverage, which is a reasonable trade in the situations it is designed for and not a general solution.
7 Does Language Structure Shape Thinking?
Underneath the measurement question sits an older one about whether the language you speak shapes how you think. It is worth addressing because it is frequently invoked to justify the question this page is about.
The strong version, that language determines the boundaries of thought, has not survived. People routinely think about distinctions their language does not encode, learn new distinctions, and translate between languages that carve up the world differently.
The weak version, that language influences habitual attention and categorisation, has genuine support in specific domains. Documented effects concern how speakers of different languages describe spatial relations, remember colour boundaries, or attend to grammatical features their language requires them to mark. These are real, they are specific, and they are modest.
None of them constitutes a difference in general cognitive ability. An effect on how attention is habitually allocated is not an effect on reasoning capacity, and no finding in this literature supports ranking languages by the cognitive capability of their speakers.
The number word effect in section 2 is the closest thing to a genuine structural effect on a cognitive measure, and it is instructive precisely because it is so narrow. It changes one subtest, through a completely understood mechanism, and it tells you nothing about the person. If the strongest documented case of language structure affecting a cognitive score turns out to be about syllable length, the general claim is not in good shape.
8 The History That Makes This Urgent
The measurement issues on this page are not academic, and the reason is a specific historical episode that shaped both the field and public suspicion of it.
In the early twentieth century, cognitive tests were administered to arriving immigrants in the United States, in English, to people who had in many cases just completed a transatlantic crossing and spoke no English at all. The reported results indicated extraordinarily high rates of what was then termed feeble-mindedness among several national groups.
The results were not treated as a measurement failure. They were reported as findings about the groups tested, and they entered public and political debate about immigration policy during a period when that debate was intensely active.
Every mechanism described on this page was operating at full force in those administrations. People were tested in a language they did not speak, on items requiring cultural knowledge they had no opportunity to acquire, under conditions of exhaustion and distress, using instruments normed on a completely different population. The scores obtained were essentially uninterpretable, and they were interpreted anyway.
Group testing conducted during the same period produced results that were similarly used in arguments about the relative capability of national groups, with the language confound treated as a minor technical caveat rather than as the dominant determinant of the scores.
The episode is why professional standards now contain explicit requirements about testing individuals whose first language differs from the language of assessment, and why any competent report documents language history. It is also a substantial part of why cognitive testing carries public suspicion that testing of other kinds does not. The suspicion is not baseless, it is a memory of a period when exactly the errors described here were made at scale and used to support policy.
Nothing about the present makes a repeat impossible. The mechanisms are the same, the temptation to read a group difference as a finding rather than an artefact is the same, and the confound remains invisible in the resulting numbers.
9 Writing Systems and Reading Measures
A related effect operates on reading and literacy measures rather than on cognitive subtests, and it matters because psychoeducational assessment frequently combines the two.
Writing systems differ in how consistently they map symbols to sounds. Some orthographies are transparent, where a letter reliably corresponds to one sound, making decoding largely rule-governed once the rules are learned. Others are opaque, where the same letters take different sounds in different words and a great deal must be learned item by item.
The consequence is that children learning to read in transparent orthographies reach accurate decoding considerably faster than children learning opaque ones. The difference is a property of the writing system rather than of the children, and it means that reading measures calibrated in one language cannot be transferred to another even when the underlying construct is the same.
It also changes how reading difficulty presents. In an opaque orthography, difficulty shows most clearly in decoding accuracy, since irregular words must be individually mastered. In a transparent one, decoding accuracy is achieved by nearly everybody and difficulty shows instead in reading speed and fluency. An assessment protocol designed around accuracy will therefore miss cases in a transparent-orthography population, and one designed around fluency will miss cases in an opaque-orthography population.
For anybody assessed in a language other than the one they learned to read in, this compounds the effects in section 3. Reading measures reflect the orthography of instruction, and a person literate in one writing system reading in another carries a cost that is neither a cognitive finding nor recoverable by adjustment.
The practical implication is the same as elsewhere on this page: the history matters, it must be recorded, and a report that does not mention it is not interpreting the scores it contains.
10 Societies Where Everybody Is Multilingual
Much of the discussion above implicitly assumes a monolingual default with second-language speakers as the exception. For much of the world that assumption is simply wrong, and the consequences deserve stating.
In many countries the language of schooling differs from the language of the home, and both differ from a regional lingua franca. People move between three or more languages daily, each dominant in a different domain: one for family, one for education and formal writing, one for commerce or media.
This produces a profile that no monolingual framework describes well. Somebody may have their largest vocabulary in the language of schooling while being far more fluent conversationally in the home language, and be most comfortable with technical material in a third. There is no single strongest language, and asking which one to test in has no clean answer.
The measurement implication is that verbal assessment in such populations measures a domain-specific slice of linguistic knowledge rather than general verbal ability. A vocabulary subtest in the language of schooling samples academic vocabulary well and everyday vocabulary poorly, and one in the home language does the reverse.
Instrument availability compounds it. Properly standardised adaptations exist for a limited set of languages, dominated by those with large wealthy markets, so assessment in much of the world proceeds either in a colonial or administrative language or with an instrument that was never locally normed. Both compromises are common and both are frequently invisible in the resulting report.
None of this is solvable by better statistics. It is a limit on what verbal cognitive measurement can do in multilingual settings, and the appropriate response is narrower claims rather than adjusted numbers, with greater weight on the nonverbal indices and explicit acknowledgement of what the verbal ones do and do not represent.
11 If This Affects You
For anybody being assessed in a language other than their strongest, several things are worth doing.
Say so, before testing. Language history is exactly the kind of background that determines how a profile is interpreted. A verbal to nonverbal gap means one thing in a monolingual examinee and something entirely different in a second-language one, and the examiner cannot tell which without being told.
Ask which language the assessment will use and whether an adaptation exists. Where a properly standardised local version exists, it is preferable to being tested in a second language on the original. Where it does not, that is worth knowing before rather than after.
Ask whether nonverbal measures are appropriate. For some purposes they are the right instrument, and for others their narrower construct coverage makes them insufficient. This is a clinical judgment and it is reasonable to raise it.
Read the verbal indices with the language in mind. A depressed verbal score in a second-language examinee is expected and is not a finding about reasoning ability. The nonverbal indices are the better estimate in that situation, and a good report will say so explicitly.
Do not treat a low verbal score as a ceiling. Vocabulary in a second language continues to grow with exposure over decades. Unlike most things a cognitive assessment measures, this one genuinely does change, which is why the score describes a current state rather than a fixed characteristic.
12 How This Applies to an Online Assessment
Everything above applies to any instrument, and unsupervised administration adds one specific problem worth stating.
A clinical examiner takes a language history at intake and interprets the profile in light of it. Software does not. An online battery administered to somebody in their second language produces exactly the same depressed verbal indices, with nothing in the process to record why, and no examiner to note it in a validity statement.
What can be done instead is reporting domains separately rather than collapsing them into one number. A person whose verbal index sits well below their fluid reasoning and visual spatial indices can see that pattern and interpret it themselves, which they cannot do from a composite that has already averaged it away.
ACIS reports six domains with percentiles and confidence intervals against a stated reference group. For anybody testing in a second language, the fluid reasoning, visual spatial, and quantitative indices are the less affected estimates, and the verbal index should be read as reflecting the language of administration rather than reasoning capacity.
That is a limitation of the format rather than a solution to it, and stating it is more useful than a composite presented as though language were not a factor. The standard limitations also apply: administration is unsupervised, conditions cannot be verified, and no institution is obliged to accept the result.
The same caution applies to the reference group. A percentile is meaningful only against a stated population, and if your linguistic background differs substantially from that population's, the comparison is doing something the norm sample was not built to support. That is worth holding in mind when reading any verbal index, from any instrument, that was normed on speakers whose relationship to the language differs from yours.
13 FAQ: Language and Cognitive Testing
Which language has the highest average IQ?
None, and the question is malformed. A language is not a population, languages span many countries and conditions, and no research programme measures cognitive ability by language.
Does the language I speak affect my intelligence?
No. It affects your measured score on verbal subtests, which is a different thing, and on one subtest it affects the score through a purely structural property of the vocabulary.
What is the number word effect?
Digit span depends on how much you can rehearse in a time-limited window, so languages with shorter number words produce longer measured spans.
How was that established?
By testing Welsh-English bilinguals in both of their own languages. The same people showed shorter spans in Welsh, where number words take longer to pronounce.
Why is the within-bilingual design important?
Because comparing two different groups would confound language with everything else that differs between them. The same individuals in both languages leaves only the linguistic property.
Do Chinese speakers have longer digit spans?
Measured spans are longer, and the explanation is that Chinese number words are unusually short. It is evidence about phonology rather than about working memory capacity.
How much does testing in a second language cost me?
Substantially on verbal measures, depending on when you acquired the language and how much you use it. Vocabulary and knowledge subtests are affected most.
Does it affect the whole battery equally?
No. Verbal indices fall most, working memory somewhat, and fluid reasoning and visual spatial least, since those depend on instructions rather than content.
Why is that uneven effect dangerous?
Because a large verbal to nonverbal gap can indicate a learning difficulty in a monolingual examinee. In a second-language examinee it indicates the language of testing.
Can a test just be translated?
No. Item difficulty does not survive translation, some items have no equivalent, norms must be rebuilt on a local sample, and structural equivalence must be demonstrated rather than assumed.
Why do adaptations lag original publications?
Because standardising on a representative sample of the new population is the expensive part, and it cannot be skipped without making the scores uninterpretable.
What should I ask about a non-English version?
Which adaptation was used and when it was standardised locally. That date, not the edition number, determines how current the norms behind your score are.
Do bilinguals have smaller vocabularies?
In each language, typically yes, while total vocabulary across both is comparable or larger. A vocabulary subtest samples one language and understates the whole.
Is there a bilingual executive function advantage?
Heavily contested. Influential early findings were followed by mixed replications, and systematic reviews have questioned whether the effect exists. The field has not converged.
Does that debate affect testing?
Not much. The measurement consequence, that a bilingual tested in one language produces an understated verbal score, holds regardless of how the executive question resolves.
Do nonverbal tests solve the language problem?
They help substantially. Instructions still require comprehension, format familiarity is not culturally neutral, and they omit crystallised knowledge entirely.
Is a nonverbal score the same as a full scale score?
No. It covers fluid reasoning and visual processing while omitting a substantial part of cognitive ability, so it should not be reported as equivalent.
Does language shape how people think?
The strong version has not survived. The weak version has support in specific domains like spatial description and colour memory, and none of it constitutes a difference in general ability.
What should I tell an examiner?
Your full language history: which languages, acquired when, used in which contexts. It determines how the whole profile is interpreted.
Is a low verbal score in my second language permanent?
No. Vocabulary in a second language continues growing with exposure over decades, so the score describes a current state rather than a fixed characteristic.
Which indices should I trust if tested in my weaker language?
Fluid reasoning, visual spatial, and quantitative are the less affected estimates. The verbal index reflects the language of administration as much as reasoning ability.
14 Best Next Step
Language does not determine ability and it substantially determines measurement. The clearest demonstration is a subtest whose scores move with syllable length, which is a fact about vocabulary rather than about anybody's memory.
For the domain structure that makes a language effect visible rather than averaged away, read Cognitive Domains. For the subtests most affected, read The Subtests Inside an IQ Test. For why group labels like languages and states do not define sampleable populations, read Average IQ by State. For your own domain profile, take the assessment.
The number word finding and the bilingualism debate both come from the peer reviewed literature. Adaptation requirements come from the international standards governing them.
Ellis, N.C. & Hennelly, R.A. (1980). A bilingual word-length effect: implications for intelligence testing and the relative ease of mental calculation in Welsh and English. British Journal of Psychology, 71(1), 43-51. The within-bilingual demonstration that digit span varies with number word pronunciation time.
International Test Commission. Guidelines for Translating and Adapting Tests. The standards governing adaptation, including requirements for local standardisation and demonstration of construct equivalence.
Paap, K.R. & Greenberg, Z.I. (2013). There is no coherent evidence for a bilingual advantage in executive processing. Cognitive Psychology, 66(2), 232-258. The influential challenge to the bilingual executive advantage literature.
Bialystok, E., Craik, F.I.M. & Luk, G. (2012). Bilingualism: consequences for mind and brain. Trends in Cognitive Sciences, 16(4), 240-250. The case for bilingual effects on cognition, including the smaller per-language vocabulary finding.
Baddeley, A.D., Thomson, N. & Buchanan, M. (1975). Word length and short-term memory. Journal of Verbal Learning and Verbal Behavior, 14(6), 575-589. The word length effect that explains why number word duration determines measured span.
American Psychological Association. Understanding psychological testing and assessment. What background information an examiner requires to interpret a profile, including language history.
McGrew, K.S. (2009). CHC theory and the human cognitive abilities project. Intelligence, 37(1), 1-10. The distinction between crystallised knowledge and fluid reasoning that determines which indices a language effect distorts.
Pearson (2008). WAIS-IV Technical and Interpretive Manual. Documents the digit span conditions and the verbal subtests most affected by language of administration.
Trahan, L.H., Stuebing, K.K., Fletcher, J.M. & Hiscock, M. (2014). The Flynn effect: a meta-analysis. Psychological Bulletin, 140(5), 1332-1360. Why the local standardisation date, rather than the edition number, determines how current an adaptation's norms are.
Voncken, L., Albers, C.J. & Timmerman, M.E. (2019). Improving confidence intervals for normed test scores. Behavior Research Methods. Open access. Why norms cannot transfer across populations without local standardisation.
Buros Center for Testing. Mental Measurements Yearbook. Independent evaluation of instruments including language-reduced and nonverbal batteries and what construct coverage they sacrifice.
Take the assessment
You get a profile, not a number
ACIS measures six CHC domains across 20 subtests and reports each one with its own normed score and confidence interval, so you can see where you are strong and where you are not.