Average IQ by Personality Type, and Why the Table Comes Out Flat
No peer reviewed study has published mean IQ scores for the sixteen Myers-Briggs types. What exists instead is a map from type to trait, and a meta-analytic coefficient for the one trait that relates to measured ability. Both are on this page, with the arithmetic shown.
Only one of the four MBTI letter pairs tracks a trait with an established relationship to measured cognitive ability, and that relationship is modest.
0 The Quick Answer on Average IQ by Personality Type
There is no published table of average IQ by personality type, because personality type is the wrong unit for the question, and the best evidence available puts the largest defensible gap between any two Myers-Briggs letter groups at roughly six IQ points, against a within group spread of fifteen. That is the honest answer, and it is more useful than the rankings circulating online, which have no sources at all.
The question is answerable, though, and refusing to answer it would be evasion rather than rigor. Here is the chain. Type systems can be translated into trait systems: McCrae and Costa established the mapping in 1989 in the Journal of Personality. Traits can then be checked against measured ability: Stanek and Ones published the definitive map in 2023 in PNAS, summarizing 60,690 relations between 79 personality constructs and 97 cognitive ability constructs across 3,543 meta-analyses. Only one of the four Myers-Briggs letter pairs corresponds to a trait with a real ability relationship, and even that one is modest.
The trait is Openness to Experience, which the sensing versus intuition letter tracks. Stanek and Ones report a meta-analytic correlation of .26 between global openness and general mental ability, with a 95 percent confidence interval of .25 to .27. On the effect size benchmarks those authors adopt, .20 is medium and .30 is large. So the relationship is genuine, replicated, and nowhere near strong enough to tell you what any individual scored.
The rest of this page does three things. It gives the type by type material, correctly labelled. It shows what the four letter label costs in measurement terms, using published coefficients rather than assertion. And it explains why a personality inventory cannot stand in for a measure of cognitive ability, no matter how well the descriptions fit.
.26
Meta-analytic correlation between Openness to Experience and general mental ability, 95 percent CI .25 to .27 (Stanek and Ones, 2023, PNAS 120(23)).
About 6 points
Upper bound on the Full Scale gap between the intuition group and the sensing group, calculated below from that coefficient and the published cost of splitting a trait in half.
Zero
Peer reviewed studies reporting mean IQ scores for all sixteen Myers-Briggs types. The rankings online are not drawn from one.
Producing a credible table of mean IQ for sixteen types would require something nobody has ever assembled: a large, representative sample given both a type inventory and a properly normed ability battery, with enough people in each of the sixteen cells to estimate a mean. The closest anyone has come is a single study from 1996, discussed later on this page, and it reported results by letter rather than by type.
The cell size problem alone is severe. The sixteen types are not equally common in any sampled population, and the rarest are reported at low single digit percentages. A study wanting a hundred people in its smallest cell would need thousands of participants who had completed both instruments under comparable conditions. Cognitive batteries are expensive to administer and are usually licensed to qualified users, which is exactly why so much personality research pairs self report with self report rather than self report with performance.
What fills the gap instead is a genre of content with no measurement behind it. The rankings you find are typically built from forum polls, from self reported scores on the free quizzes that sit alongside validated batteries, or from nothing traceable at all. Those three sources share a fatal property: the people who supply the numbers choose whether to supply them. Someone who scores well on a free web test and identifies with a type description has every reason to post the pairing. Someone who scores poorly does not. The result is a sample assembled by the exact variable under study, which is a separate problem from whether an unproctored test can be accurate at all.
This matters more than it might seem, because a self selected sample does not merely add noise. It can invert a relationship. If the members of one type are more likely to be recruited into online intelligence discussions, that type will appear more able even if the underlying trait difference is zero. Understanding how a reference sample is built is what separates a norm from a leaderboard. A claim can also outrun its measurement entirely rather than merely bias it. The widely repeated finding that generative artificial intelligence lowers intelligence rests on three studies that recorded EEG connectivity, survey answers and perceived mental effort, so the quantity in that headline was never administered to anyone.
So the responsible move is not to invent the table. It is to answer the question through constructs that have been measured properly, and to be explicit about every step of the translation.
What is not being claimedNothing on this page reports an ACIS statistic for any personality type. ACIS collects no personality data of any kind, has never administered a type inventory, and publishes no relationship between type and score. Every personality coefficient here comes from the external literature and is attributed in the text.
2 What the MBTI Is, and What It Says About Itself
The Myers-Briggs Type Indicator is popular for a reason that has nothing to do with intelligence: it produces a rich, engaging, personally resonant description from a short questionnaire, and it does capture real variation in how people describe themselves. Dismissing it as a horoscope misses what it does and what it is criticized for.
The instrument descends from Carl Jung's account of psychological types, extended by Katharine Cook Briggs and Isabel Briggs Myers, who added the fourth dimension of judging versus perceiving. A respondent answers forced choice items and receives four letters, one per dimension. Randy Stein and Alexander Swan, in a 2019 review in Social and Personality Psychology Compass, report that the instrument is taken by roughly two million people a year and is widely used in corporate development. That reach is a fact about the world, not an argument for or against the instrument.
Two things about it are directly relevant here, and both come from the MBTI's own framework rather than from its critics.
First, the theory holds that the letters report preferences, not tendencies and not abilities. Stein and Swan note this explicitly as a defining feature of the theory. The publisher has made the same point in defending the instrument against performance based criticism, arguing that it was never intended to predict outcomes. Taken at face value, that position already answers the search query: an instrument that disclaims measuring abilities cannot rank types by ability, and the rankings that circulate are not the publisher's claim.
Second, the theory asks the letters to do something the underlying scores cannot support. Nothing about the item content produces two clumps of people. The scores come out as a continuous distribution, and the letters are made by cutting that distribution somewhere near the middle. Whether that cut is defensible is the central technical question, and it has a quantitative answer that appears later on this page.
Between those two points sits the fair verdict. The MBTI is a reasonable way to open a conversation about how people differ. It is not a substitute for a cognitive ability test, and it does not claim to be.
Fair criticism, precisely statedThe complaint against the MBTI is not that its dimensions are fake. Three of them map onto well established personality factors. The complaint is that it dichotomizes continuous distributions and reports categories rather than positions, and that the categories then get treated as kinds of people.
3 Translating Type Talk Into Trait Talk
The bridge that makes this question answerable at all was built by Robert McCrae and Paul Costa in 1989, in a paper titled Reinterpreting the Myers-Briggs Type Indicator From the Perspective of the Five-Factor Model of Personality, published in the Journal of Personality, volume 57, issue 1, pages 17 to 40. Their data came from 267 men and 201 women aged 19 to 93, using both self reports and peer ratings on the NEO Personality Inventory.
Two findings from that paper carry the argument. The first is negative: the authors report no support for the view that the instrument measures truly dichotomous preferences or qualitatively distinct types, concluding instead that it measures four relatively independent dimensions. The second is constructive: the four indices did measure aspects of four of the five major dimensions of normal personality, which means type language can be translated into trait language without loss of meaning.
That translation is what licenses everything that follows. Once each letter pair is understood as a position on a familiar trait, the enormous literature relating those traits to measured ability becomes available. And once it becomes available, the answer to the search query turns out to be lopsided.
Note what the mapping leaves out. There is no Myers-Briggs letter corresponding to the fifth factor, emotional stability. That omission matters because facets of neuroticism carry some of the more substantial ability relationships in the literature, a point developed below. A type label is therefore silent on the trait dimension where several of the sizable negative coefficients live.
Letter pair
Five factor dimension it tracks
That dimension with general mental ability
What the letter licenses
Extraversion versus introversion
Extraversion
-.02, negligible
Nothing about measured ability
Sensing versus intuition
Openness to Experience
.26, medium
The only real signal of the four
Thinking versus feeling
Agreeableness, with feeling as the high pole
-.01, negligible
Nothing about measured ability
Judging versus perceiving
Conscientiousness, with judging as the high pole
.01 for the factor; a judging versus perceiving compound trait sits at -.11
Small at most, and direction dependent on scoring
Mapping from McCrae and Costa (1989). Coefficients from Stanek and Ones (2023). The compound trait row is reported by Stanek and Ones among the compound traits they group with conscientiousness, and its magnitude of about a tenth sits between very small and small on the benchmarks they adopt.
4 Openness Is the Only Letter With a Real Ability Link
The single most important number for this question comes from Kevin Stanek and Deniz Ones, writing in the Proceedings of the National Academy of Sciences, volume 120, issue 23, in 2023. Their paper, Meta-analytic relations between personality and cognitive ability, quantifies 60,690 relations between 79 personality constructs and 97 cognitive ability constructs across 3,543 meta-analyses drawn from a century of research.
Global openness correlates .26 with general mental ability, the g factor those studies are estimating, with a 95 percent confidence interval of .25 to .27. That interval is narrow because the underlying evidence is vast, which is precisely why this coefficient is worth more than any single study. For comparison within the same paper: conscientiousness sits at .01, extraversion at -.02, agreeableness at -.01, and neuroticism at -.08.
The authors are explicit that substantial relations rarely appear at the Big Five factor level and instead emerge at the level of aspects, facets and compound traits. Within openness, the split is stark. The intellect aspect, covering engagement with ideas and complex problems, correlates .28 with general mental ability on average, with the ideas facet reaching .40 in their summary table. The experiencing aspect, covering aesthetics and fantasy, is negligible with most non invested abilities and sits at .01 with fluid abilities. Curiosity about ideas and curiosity about art are not the same variable, which is also why the creativity literature separates from the ability literature as sharply as it does.
That distinction is the whole reason the intuition letter carries any signal at all. It also explains why the signal is smaller than people expect: the letter blends both aspects into one cut.
Where the relationship shows up is just as informative as how large it is. Openness relates far more strongly to acquired verbal knowledge than to reasoning under novelty, and barely at all to acquired quantitative ability. That pattern maps directly onto the six domains a modern battery reports separately, which is why a single composite figure hides the most interesting part.
Here is the type by type material, built the only way it can honestly be built: by projecting each four letter combination onto the traits it tracks, and then applying the one trait coefficient that relates to measured ability. The final column is an upper bound derived in a later section, not an observed mean, and it is deliberately identical for every type sharing the same second letter.
Read the flatness as the finding. Eight types sit at most three points above the adult mean and eight sit at most three points below it, and the word governing that sentence is at most. The bound assumes the intuition letter is a perfect reading of openness, which it is not. The true separation is smaller than what the table shows.
Type
Trait profile implied by the four letters
Openness pole
Upper bound on the Full Scale difference from the adult mean
INTJ
Low extraversion, high openness, low agreeableness, high conscientiousness
High
At most about 3 points above
INTP
Low extraversion, high openness, low agreeableness, low conscientiousness
High
At most about 3 points above
INFJ
Low extraversion, high openness, high agreeableness, high conscientiousness
High
At most about 3 points above
INFP
Low extraversion, high openness, high agreeableness, low conscientiousness
High
At most about 3 points above
ENTJ
High extraversion, high openness, low agreeableness, high conscientiousness
High
At most about 3 points above
ENTP
High extraversion, high openness, low agreeableness, low conscientiousness
High
At most about 3 points above
ENFJ
High extraversion, high openness, high agreeableness, high conscientiousness
High
At most about 3 points above
ENFP
High extraversion, high openness, high agreeableness, low conscientiousness
High
At most about 3 points above
ISTJ
Low extraversion, low openness, low agreeableness, high conscientiousness
Low extraversion, low openness, high agreeableness, high conscientiousness
Low
At most about 3 points below
ISFP
Low extraversion, low openness, high agreeableness, low conscientiousness
Low
At most about 3 points below
ESTJ
High extraversion, low openness, low agreeableness, high conscientiousness
Low
At most about 3 points below
ESTP
High extraversion, low openness, low agreeableness, low conscientiousness
Low
At most about 3 points below
ESFJ
High extraversion, low openness, high agreeableness, high conscientiousness
Low
At most about 3 points below
ESFP
High extraversion, low openness, high agreeableness, low conscientiousness
Low
At most about 3 points below
Trait projection from the McCrae and Costa (1989) mapping. The bound is calculated from the Stanek and Ones (2023) openness coefficient and the dichotomization factor published by MacCallum and colleagues in 2002, with the arithmetic shown further down. No cell in the last column is an observed group mean, and none should be quoted as one.
One further caution about reading this table backwards. A three point bound describes a difference between group averages of thousands of people. It says nothing about you. Two individuals drawn from the two groups will differ by whatever they differ by, and the overlap between the distributions is enormous. This is the same reasoning that governs group averages by profession and group averages by college major, where the between group spread is likewise dwarfed by the spread inside each group.
6 The One Study That Actually Measured Both
Alan Kaufman, James McLean and Alan Lincoln published the closest thing to a direct test of this question in 1996, in the journal Assessment, volume 3, pages 225 to 239, under the title The Relationship of the Myers-Briggs Type Indicator to IQ Level and the Fluid and Crystallized IQ Discrepancy on the Kaufman Adolescent and Adult Intelligence Test. It remains the reference point, and it is almost never cited by the pages ranking types by intelligence.
The design was unusually good for this literature. Participants numbered 1,297, aged 14 to 94, tested across the United States during the nationwide standardization of the KAIT. That means the ability measure was a properly administered battery with a real normative programme behind it, not a web quiz, and the participants were recruited for norming rather than for a study about personality.
The authors hypothesized that people favoring intuition and thinking would score higher and would show an advantage for fluid over crystallized ability. The first half held in one place only: people classified as intuitive earned higher KAIT Composite IQs than people classified as sensing. That is a real, independently obtained result in the same direction as the openness coefficient, which is worth stating plainly.
The second half failed. The fluid versus crystallized discrepancy was not meaningfully related to any Myers-Briggs dimension. Type predicted, at best, a small difference in overall level, and predicted nothing about the shape of the cognitive profile.
That failure is more interesting than the success, because profile shape is where the practically useful information lives. Knowing someone scores a little above average tells you almost nothing actionable, even for the outcome ability predicts best, which is performance in formal education. Knowing that their reasoning under novelty runs well ahead of their acquired knowledge, or the reverse, tells you a great deal, and it is the reason the fluid and crystallized distinction exists in the first place. A type label cannot supply it.
How to label this resultThe Kaufman finding is documented: it comes from a peer reviewed study with a named sample and a standardized instrument. It should not be inflated into a ranking of types, because the study reported by letter and not by type, and it reported one significant contrast out of a larger set of hypotheses.
7 What the Cut Costs, in Published Numbers
The technical objection to type reporting is not a matter of taste, and it has been quantified. Robert MacCallum, Shaobo Zhang, Kristopher Preacher and Derek Rucker set out the arithmetic in On the Practice of Dichotomization of Quantitative Variables, published in Psychological Methods, volume 7, issue 1, pages 19 to 40, in 2002.
When a normally distributed continuous variable is split at its mean, the correlation it can show with anything else is multiplied by a constant the authors designate d, which equals .798 at that split point. Shared variance falls to .637 of its original value. The loss of statistical power is equivalent, for a two tailed test at the conventional alpha, to discarding close to 36 percent of the sample. And in a later section of the same paper the authors show that the reliability index of the measure is reduced by exactly the same factor.
That last consequence is the one that bites hardest for type reporting. The information thrown away is not noise. It is the part of the score that distinguishes a person sitting just past the cut from a person sitting far past it, and that is precisely the information a reader wants when they ask which type is smartest. The authors state that they know of no findings of positive consequences of dichotomization.
Apply the factor to the case at hand. Openness relates to general mental ability at .26. Split openness in half at the mean and the best available correlation between the resulting two groups and measured ability becomes .798 times .26, which is about .21. That is the ceiling for what any sensing versus intuition contrast can carry, before any allowance for the fact that a letter is not a perfect reading of the trait.
This is a general lesson about measurement rather than a complaint about one instrument. It is the same reason a good battery reports scaled scores rather than pass and fail verdicts, and the same reason a score is interpreted as a position with an error band rather than as a category. Categories are cheap to communicate and expensive in information. The cost surfaces wherever a self report is given a threshold: aphantasia is defined by a cut on a sixteen item imagery questionnaire, and moving that cut swings the published prevalence from 0.7 percent to nearly 5 percent of the population, which is the same boundary problem in a different self report.
.798
The factor by which a correlation shrinks when a continuous variable is split at its mean (MacCallum, Zhang, Preacher and Rucker, 2002).
.637
The proportion of the original shared variance that survives the same split, in the same paper.
About 36 percent
The effective loss of sample size from a median split, expressed as the sample you would have to discard to lose the same statistical power.
8 Why Your Letters Move Between Administrations
The instrument's continuous scores are reasonably stable. Its four letter label is not, and the gap between those two statements is the whole problem. Ken Randall, Mary Isaacson and Carrie Ciro established the first half in a systematic review and meta-analysis published in 2017 in the Journal of Best Practices in Health Professions Diversity, volume 10, issue 1, pages 1 to 27, available as the full text of that review.
They screened 221 candidate studies and included seven. Three examined test retest reliability and could be pooled, giving 314 participants across a mean interval of 13.93 months. The pooled random effects coefficients were .764 for extraversion and introversion, .753 for sensing and intuition, .612 for thinking and feeling, and .775 for judging and perceiving. Their own conclusion, stated cautiously, is that the instrument performs reliably over time. On the continuous scores, that is a defensible reading.
Note the limits the authors themselves flag. Heterogeneity across the pooled studies was substantial, all participants were college age, only older forms of the instrument were represented, no included study examined internal consistency, and study quality scores ran from 9 to 14 out of a possible 20. Those limits are stated in the paper and should travel with the coefficients.
Now apply the second half. A retest correlation of .753 on a continuous score does not mean the label built from that score is 75 percent stable. Where you land relative to the cut depends on how far your score sits from it. Treating the two administrations as bivariate normal with that correlation and the cut at the mean produces the following, which is a calculation from the published coefficient rather than an observed rate.
Where your continuous score sits
Probability the letter flips on retest
Exactly at the cut
50 percent
Half a standard deviation past the cut
About 28 percent
One standard deviation past the cut
About 13 percent
Two standard deviations past the cut
About 1 percent
The point is not that the letter is random. For someone far from the cut it is quite stable. The point is that stability is a property of the person's position rather than of the label, and the label does not report the position. Two people can both be told they are intuitive when one of them will be told the opposite next year. This is the difference between reporting a category and reporting a score with a stated standard error of measurement.
9 The Sixteen Types Do Not Show Up in the Data
If sixteen distinct kinds of people existed, the scores would show it, and they do not. The prediction is straightforward: a real dichotomy should produce two clusters of scores on each dimension, with a thin region between them. What the distributions actually show is a single hump with most people near the middle, which is exactly where the cut falls.
For a while a counterargument circulated, based on reports that the scores were bimodal after all when computed with item response theory. Timothy Bess and Robert Harvey settled it in the Journal of Personality Assessment, volume 78, issue 1, pages 176 to 186, in 2002, in a paper whose subtitle asks whether the bimodality is fact or artifact. Working with roughly 12,000 people from leadership development programmes, they showed that the earlier bimodality reports were artifacts of the scoring software's default setting for the number of quadrature points. Raising that number from ten to fifty produced distributions the authors describe as strongly center weighted.
A second line of evidence attacks the same question from a different angle. Statistical taxometrics asks directly whether a set of scores comes from two latent classes or from one continuum. Stein and Swan, in their 2019 review, report that Arnau and colleagues applied Paul Meehl's coherent cut kinetics methods to several Jungian type assessments including this one in 2003, and found no support for the existence of underlying types in any of them.
So three independent lines converge. McCrae and Costa found no support for truly dichotomous preferences in 1989. Bess and Harvey removed the bimodality evidence in 2002. Taxometric methods found no latent classes in 2003. The dimensions are real; the sixteen boxes are a reporting convention laid over them.
This is worth separating from a cheap dismissal. A reporting convention can still be useful for conversation, and the fact that a boundary is drawn rather than discovered does not make the underlying variation imaginary. It does mean that a question phrased as which box is smartest has no measured answer, and that the wider debate about categorizing minds keeps running into the same wall.
Continuous does not mean meaninglessHeight is continuous, and tall is still a useful word. The objection to type reporting is specific: it is that the label is presented as an identity that a person has, rather than as a shorthand for a position they occupy, and that the presentation encourages exactly the inference this page is about.
10 How Large Could the Gap Be, at Most
Putting the published coefficients together produces a defensible ceiling, and the arithmetic is short enough to check. What follows is a calculation from two external published figures, not a measurement, and every input is named.
Start with the openness coefficient of .26 from Stanek and Ones. Apply the dichotomization factor of .798 from MacCallum and colleagues, since the sensing and intuition letter is a split of that trait. The resulting point biserial correlation between group membership and general mental ability is about .21. Converting a point biserial of .21 with two equally sized groups into a standardized mean difference gives roughly 0.42 standard deviations. On a scale where the standard deviation is fifteen points, that is about six and a third IQ points between the two group means.
Three properties of that number deserve emphasis, because a bare six can be misread as large.
It is a ceiling, not an estimate. It assumes the letter is a flawless reading of openness, which it is not, since the intuition scale blends the intellect aspect that carries the signal with the experiencing aspect that does not. Every departure from that assumption shrinks the number, and the real separation is smaller.
It describes almost nothing about individuals. A standardized difference of 0.42 means that if you pick one person at random from each group, the person from the intuition group scores higher about 62 percent of the time. A coin flip is 50 percent. The two distributions overlap by roughly 83 percent of their area. Anyone using a type label to guess a person's score is right slightly more often than chance and wrong very frequently.
And it is small relative to the precision of a decent measurement. The standard error of measurement for the ACIS Full Scale IQ is about 1.60 IQ points, drawn from a technical analysis set of 2,750 complete records with a self selected, unsupervised sample and a modelled adult reference frame. The entire best case difference between the eight intuition types and the eight sensing types is roughly four times that error band. It is real, and it is far too small to do what a ranking asks of it.
0.42 SD
Standardized difference implied at the ceiling, equal to about six and a third IQ points between the two letter groups.
62 percent
Chance that a randomly chosen person from the higher group outscores a randomly chosen person from the lower one, at that ceiling.
About 83 percent
Overlap between the two distributions at that same ceiling, which is what makes individual prediction hopeless.
Two types dominate every informal ranking, and the pattern is explainable without any of them being smarter. Four mechanisms are enough, and each is independently documented.
The first is the one real coefficient. The intuition letter tracks openness, and openness relates to ability at .26. Every ranking that puts N types on top is picking up a genuine signal, then inflating it by an order of magnitude. Kaufman and colleagues found the same direction with a real battery in 1996. The direction is right and the magnitude is fantasy.
The second is the shape of the openness relationship. It is strongest with acquired verbal abilities at .29 and weakest with acquired quantitative abilities at .07, in the Stanek and Ones figures. Verbal fluency is the most visible form of intelligence in a text based forum. Types that project high openness will therefore read as smart in exactly the settings where these rankings are written, whatever the rest of their profile looks like.
The third is selection into the conversation. People who enjoy introspective self analysis are the ones who take type inventories, post their results, and discuss intelligence online. Introversion and openness both push in that direction. The sample of people supplying data about their own type and their own score is assembled by the traits under study, which is a textbook route to a spurious effect.
The fourth is the description itself. The type descriptions for the N types lean on words like analytical, strategic, theoretical and inventive. Those words are compliments about thinking. A person who reads a flattering description of their own reasoning and then reports that their type is intelligent has reported a fact about the description, not about themselves. The comparison ACIS refuses to make elsewhere applies here too: behavioural signs read as intelligence are not measurements of it. The self report underneath the rankings is weak in a way that has been quantified, since pooled across 41 published studies a person's estimate of their own intelligence correlates about .33 with their measured score, leaving a prediction interval roughly 55 points wide, which is what a guess at one's own standing is actually worth.
Stack the four and you get a large apparent effect built on one modest real one, assembled by the same process that keeps the durable myths about testing in circulation. The correction is not to declare the whole area meaningless. It is to notice that the question changed somewhere along the way, from what does openness predict to which box is best, and only the first version has an answer.
A note on the T in INTPThe thinking letter is the one people most often read as an ability claim, and it is the one with the least support. It tracks the low pole of agreeableness, which correlates -.01 with general mental ability in the Stanek and Ones synthesis. Preferring impersonal criteria is not a cognitive advantage.
12 What to Do If You Actually Want the Number
If the question behind the search is where do I actually stand, the instrument that answers it is an ability measure, and the useful output is a profile rather than a single figure. That is not a marketing claim; it follows directly from the evidence above.
Recall the pattern in the openness coefficients. The trait relates to acquired verbal ability at .29, to fluid reasoning at .19, to working memory at .14, to visual spatial processing at .10, to processing speed at .08, and to acquired quantitative ability at .07. A person high in openness is therefore likely to be uneven across domains, not uniformly elevated. A single composite averages that unevenness away, and a four letter label never had it to begin with.
ACIS reports twenty subtests across six primary cognitive domains, with scaled scores on a mean of 10 and a standard deviation of 3, and composites on a mean of 100 and a standard deviation of 15. The Full Scale IQ composite has an omega of .9886 and a g loading of .958, with the standard error of measurement of about 1.60 points noted earlier. The higher order g model fits at CFI .9761, TLI .9726, RMSEA .0406 and SRMR .0217, with a chi-square of 916.703 on 166 degrees of freedom. Those figures come from the technical analysis set of 2,750 complete records, within an adult reference frame of 3,243 English speaking records aged 16 to 90, and they are set out in full in the published technical documentation.
Every one of those figures carries the same limitation, and it is not a small one. Administration is unsupervised, the sample is self selected rather than census based, and the adult reference frame is modelled rather than drawn from a population register. Good internal structure does not convert an online battery into a supervised clinical assessment, and this one is not a substitute for supervised testing when a formal decision depends on the result.
For researchers who want type and ability data collected together properly, the missing study described earlier is still missing, and it is not hard to design. It needs both instruments, a sample that was not recruited through either construct, and enough participants for the smallest cells.
13 Standards, Limits, and What This Page Does Not Claim
Every claim on this page is either sourced to a named publication or shown as arithmetic performed on sourced figures, and the two categories are kept separate on purpose. The sourced coefficients are the ones from Stanek and Ones 2023, McCrae and Costa 1989, Kaufman, McLean and Lincoln 1996, Randall, Isaacson and Ciro 2017, MacCallum and colleagues 2002, Bess and Harvey 2002, and Stein and Swan 2019. The upper bound of about six points and the letter flip probabilities are calculations, labelled as such where they appear.
ACIS has never administered a personality inventory, collects no personality data, and reports no relationship between type and score. Nothing here should be read as an ACIS finding about any type. The ACIS psychometric figures quoted are limited to the published technical values, and they are always accompanied by their limitation: unsupervised administration, a self selected sample, and a modelled adult reference frame rather than a census based one.
The instrument itself is bounded in the same way as everything else on this site. It is not a clinical or diagnostic instrument. It is not appropriate for diagnosis, for hiring decisions, for accommodation requests, or for admission to a high IQ society, and readers weighing the requirements those societies publish should note that unsupervised results are generally not accepted. It reports cognitive performance across fluid reasoning, comprehension knowledge, quantitative reasoning, visual spatial processing, working memory and processing speed, and it reports nothing about character, motivation, values or type.
That boundary is not an ACIS invention. The Standards for Educational and Psychological Testing, published jointly in 2014 by the American Educational Research Association, the American Psychological Association and the National Council on Measurement in Education, require that a test's intended interpretations be specified and that evidence be presented for each of them. A score interpreted as a personality type, or a type interpreted as a score, is an unsupported interpretation under that framework, and the APA testing standards materials make the same point about staying inside the evidence a test has actually produced. The International Test Commission guidelines on test use add the practical corollary: the person administering a test is responsible for confining its use to the purposes its validity evidence supports.
Applied to the search that brought you here, the standards give a clean instruction. Ask what construct you want measured, choose an instrument with published evidence for that construct, and read the result as a position with an error band. That procedure answers the question. Ranking sixteen labels does not.
14 Frequently Asked Questions
What is the average IQ of each MBTI personality type?
No peer reviewed study has published mean IQ scores for the sixteen types. The rankings online are built from forum polls and self reported web test scores, both of which are assembled by the people who choose to report, which makes them unusable as averages.
Which personality type has the highest IQ?
No type does, in any measured sense. The eight types carrying the intuition letter project onto higher Openness, and Openness relates to general mental ability at .26 in Stanek and Ones 2023, but that produces a group gap of at most about six points with enormous overlap.
Is INTP or INTJ smarter?
Nothing in the published evidence separates them. They differ only in the fourth letter, which tracks Conscientiousness, and Conscientiousness correlates .01 with general mental ability in the Stanek and Ones synthesis. Any ranking placing one above the other is reporting a preference, not a measurement.
Do intuitive types really score higher than sensing types?
Slightly, in one well designed study. Kaufman, McLean and Lincoln found in 1996 that people classified as intuitive earned higher KAIT Composite IQs than those classified as sensing, in a sample of 1,297 adults from the test's standardization programme.
Why does the MBTI correlate with intelligence at all?
Through one letter only. The sensing versus intuition index tracks Openness to Experience, which McCrae and Costa established in 1989, and Openness is the one Big Five factor with a substantial relationship to measured ability.
What is the correlation between Openness and IQ?
Stanek and Ones report .26 with general mental ability, with a 95 percent confidence interval of .25 to .27, drawn from 3,543 meta-analyses. On the effect size benchmarks they use, .20 counts as medium and .30 as large.
Is the MBTI a scientifically valid test?
Its dimensions map onto real personality variation, and its continuous scores show acceptable retest reliability. The specific criticism is narrower: it splits continuous distributions into categories and reports a type rather than a position, which discards information.
How reliable is the MBTI over time?
Randall, Isaacson and Ciro pooled three studies in 2017 and reported retest coefficients of .764, .753, .612 and .775 for the four scales over a mean interval of about fourteen months. Those figures describe the continuous scores, not the four letter label.
Why does my MBTI type keep changing?
Because your letter depends on which side of a cut your score falls, and scores move a little between administrations. Someone sitting right at a cut has a coin flip chance of switching, while someone two standard deviations away almost never does.
What does dichotomization mean and why is it a problem?
It means converting a continuous score into two categories. MacCallum and colleagues showed in 2002 that splitting at the mean multiplies any correlation by .798, cuts shared variance to .637 of its value, and costs the equivalent of discarding 36 percent of a sample.
Are MBTI score distributions bimodal?
No. Bess and Harvey examined roughly 12,000 respondents in 2002 and found that earlier bimodality reports were artifacts of a default setting in the scoring software. With better settings the distributions were strongly weighted toward the center.
Does the Big Five predict IQ better than the MBTI?
It predicts more transparently, because it reports positions rather than categories. The gain is not a stronger relationship but a preserved one, since no cut is applied and no information is thrown away before the correlation is computed.
Which Big Five trait relates most to cognitive ability?
Openness, and specifically its intellect side. Stanek and Ones report the intellect aspect at about .28 with general mental ability and the ideas facet at .40, while the experiencing aspect covering aesthetics and fantasy is close to zero.
Does conscientiousness relate to intelligence?
The broad factor sits at .01, which is nothing. Its industriousness aspect is a different story at .32 with general mental ability, which illustrates the paper's central message that the action is below the factor level.
Are introverts smarter than extraverts?
No. Extraversion correlates -.02 with general mental ability in the Stanek and Ones synthesis, which is indistinguishable from zero. The stereotype survives because quiet, bookish behaviour is easy to read as intelligence.
Can a personality test estimate my IQ?
No. Even at the most generous reading, a letter based split of Openness picks the higher scorer out of a random pair only about 62 percent of the time. That is barely better than a coin flip and useless for an individual estimate.
Does the MBTI claim to measure ability?
It explicitly does not. The framework holds that the letters report preferences rather than abilities or tendencies, and the publisher has defended the instrument on the grounds that it was never meant to predict performance.
Do the sixteen types exist as real categories?
The evidence says no. McCrae and Costa found no support for genuinely dichotomous preferences in 1989, and taxometric analyses reported by Stein and Swan found no latent classes underlying Jungian type assessments.
Should I use type to guess someone's intelligence?
No, and the overlap figure explains why. At the most favourable estimate the two letter groups overlap by roughly 83 percent of their distributions, so a type label tells you almost nothing about the person in front of you.
Has ACIS studied personality type and scores?
No. ACIS collects no personality data, has never administered a type inventory, and publishes no relationship between type and cognitive score. Every personality figure on this page comes from external published research and is attributed there.
What should I take if I want an actual cognitive score?
An ability battery that reports domain level results, since the interesting variation is between domains rather than in one number. Remember that any unsupervised online result, including this one, is not a clinical or diagnostic assessment.
Take the assessment
You get a profile, not a number
ACIS measures six CHC domains across 20 subtests and reports each one with its own normed score and confidence interval, so you can see where you are strong and where you are not.