Average IQ by College Major What the Rankings Are Really Built From
Nobody has IQ tested a physics department. Every major ranking traces back to graduate admission test scores, and the ordering is stable, real, and about who applies rather than what studying does.
0 Quick Answer
Updated August 16, 2026 by Structural. There is no dataset of IQ scores by academic major. Every ranking in circulation is derived from graduate admission test performance, primarily the GRE, whose results by intended field of study are published annually by the testing organisation.
Direct answer: the ordering is consistent across decades and across the verbal and quantitative sections in a way that is worth stating precisely. Physics and mathematics dominate quantitative performance. Philosophy consistently produces the highest verbal scores of any field, frequently the highest overall combination. Economics performs strongly on both. Education and several applied professional fields sit at the lower end of both.
What this ordering reflects is overwhelmingly selection rather than effect. Fields with demanding mathematical or analytical entry requirements attract and retain students who were already strong on those dimensions, and the admission test measures the result of that filtering. Studying philosophy does not produce a high verbal score, and people with high verbal scores are disproportionately drawn to philosophy.
Two sources feed essentially every major ranking, and knowing which one a given table used explains most of the discrepancies between them.
Graduate admission test data. The Educational Testing Service publishes annual summary statistics for GRE General Test takers, broken down by intended graduate field of study. These reports give mean verbal, quantitative, and analytical writing scores per field, along with the number of test takers. The data is public, methodologically transparent, and collected consistently, which is why it is the backbone of every serious version of this question.
The crucial detail is the phrase intended field of study. These are people planning to enter graduate study in a field, self-reported at the time of testing, and typically holding an undergraduate degree in or near it. So the data describes prospective graduate students rather than everybody who studied the subject as an undergraduate.
Undergraduate admission test data. SAT and ACT scores by intended major have been published at various points, describing school leavers stating a preference rather than anybody who completed a degree. This population is much less filtered and consequently produces a compressed ordering with the same general shape.
Neither source is an IQ measure. Both are admission tests designed to predict academic performance, built around verbal reasoning, quantitative reasoning, and writing. They correlate substantially with general cognitive ability, which is why the conversion is attempted, and they are not interchangeable with it for the reasons in section 6.
Anybody encountering a table of IQ figures by major should ask which of these two sources it came from and what transformation was applied. In most cases the answer is the GRE data with an undisclosed conversion, and the underlying ordering is identical to the published score ordering.
2 The Stable Ordering
The pattern has held across decades of data and across changes to the test itself, which is the strongest thing that can be said for it.
English literature, history, classics, linguistics
Balanced and near the mean
Middle
Biology, psychology, political science, anthropology
Below the mean on both
Lower end
Education, social work, several applied professional fields
Two features of this table deserve attention before anything is concluded from it.
The first is that the split between verbal and quantitative performance is as informative as the overall ordering and gets discarded when the data is converted to a single IQ figure. A physics cohort and a classics cohort can produce similar combined scores through entirely different profiles, and collapsing them into one number destroys the most interesting thing the data contains. This is the same information loss that makes a single composite less useful than a domain profile, discussed in Cognitive Domains.
The second is that the differences between adjacent fields are frequently small relative to the variation within any single field. A ranking presents the ordering as though it were sharp, and the underlying distributions overlap heavily. The best students in a lower ranked field outperform the median student in a higher ranked one comfortably and routinely.
3 Why Philosophy Tops the Verbal Ranking
The consistent appearance of philosophy at or near the top surprises people who expected the sciences to dominate everything, and the explanation is instructive about what the data measures.
Philosophy at undergraduate level is unusual in that it offers no vocational path and no obvious credential value. Students who choose it are, on average, choosing it because they find the material compelling rather than because it leads somewhere, and that self-selection is unusually strong compared with fields that carry an employment expectation.
The subject matter is also close to what a verbal reasoning test measures. Analysing arguments, identifying hidden premises, tracking precise distinctions between similar terms, and evaluating whether a conclusion follows are the daily activity of philosophical training and the explicit content of verbal reasoning items. Few fields have a closer match between their disciplinary practice and a test section.
Philosophy graduates also perform strongly on the quantitative section relative to other humanities fields, which follows from the same alignment. Formal logic is a standard component of philosophical training, and the reasoning it develops is structurally similar to quantitative reasoning even without the arithmetic content.
The filtering is severe at the graduate stage as well. Doctoral programmes in philosophy are among the most competitive relative to their size, with very few positions, which means the population of prospective graduate students has been through unusually heavy selection before ever taking the test.
The general lesson is that the ranking rewards fields whose disciplinary content resembles what the test measures and whose students chose them without external incentive. That is a statement about the interaction between a field and a specific instrument, not a statement about the intellect of anybody who studied it.
4 Selection Versus Effect
The question underneath all of this is whether demanding fields produce capable people or attract them. The evidence supports the second far more strongly than the first, and understanding why matters for anybody weighing what to study.
Selection operates at several stages, and each one filters. Prerequisite requirements exclude students who did not complete earlier mathematics or language courses. Introductory courses in demanding fields have high failure rates that redirect students elsewhere. Progression through a degree filters again. And the decision to apply for graduate study filters hardest of all, since it selects the subset who performed well enough to consider continuing.
By the time somebody appears in GRE data as a prospective physics graduate student, they have passed through four or five successive filters, each correlated with quantitative ability. The resulting mean says a great deal about the filters and comparatively little about physics education.
Effects do exist and are smaller than the selection. Education raises cognitive test performance measurably, with a meta-analysis of quasi-experimental studies estimating gains of roughly one to five IQ points per additional year of schooling. That is a real causal effect established through designs that separate education from selection. It is also an effect of schooling in general rather than of a particular major, and it is far smaller than the between-field differences in the rankings.
Field-specific training does produce field-specific improvement. Studying mathematics improves mathematical reasoning, and studying philosophy improves argument analysis, both of which show up on the corresponding test sections. Whether this constitutes an increase in general ability or a well-practised specific skill is precisely the transfer question examined in Can You Improve Your IQ?, and the answer depends on how closely the training matches the ability in question.
The practical upshot is that the ranking mostly describes who ends up where. It does not predict what an individual will gain from a course of study, and choosing a major to raise your score would be an unusually inefficient way to pursue that goal.
5 The Range Restriction Problem
This is the technical issue that limits what the data supports, and it is the reason field means cannot be read as population statements.
The people in GRE data are not a sample of anything general. They completed secondary education, entered university, completed or nearly completed a degree, decided to pursue graduate study, and paid to sit an admission test. Each step selects on characteristics correlated with cognitive ability.
The consequence is that the whole population under analysis sits well above the general mean and, crucially, has a compressed variance. When a variable's range is restricted, correlations involving it shrink and differences between subgroups are measured on a truncated scale. Field means computed within this population cannot be extrapolated to everybody who studied those fields, let alone to everybody who works in them.
The restriction is also uneven across fields, which is the part that undermines direct comparison. Fields where graduate study is the standard path produce test-taker populations close to their whole graduate cohort. Fields where most graduates enter employment directly produce test-taker populations consisting only of the minority choosing to continue, which is a much more selected group. Two fields with identical undergraduate populations will therefore show different GRE means purely because different fractions of them sit the test.
A concrete illustration: a field where nearly everybody proceeds to graduate school contributes its full ability distribution to the data. A field where a small percentage proceeds contributes only its upper tail. The second field will appear stronger in the rankings, and the appearance is an artefact of who took the test.
Education is the clearest illustration of this mechanism, and it is the field these rankings treat most unfairly. A graduate qualification is a standard and frequently required step in teaching, so a very large fraction of education graduates sit the test, including many already working and returning for certification. Compare that with a field where postgraduate study is unusual and only the strongest few percent continue. The education population contributes something close to its whole distribution while the comparison field contributes its upper tail, and the resulting gap in the table is substantially an artefact of that difference rather than a finding about the people.
Nobody publishing these rankings adjusts for this, because the adjustment would require knowing the ability distribution of the non-testing majority in each field, which is exactly the unavailable information the whole exercise is trying to substitute for.
6 Why the Conversion to IQ Fails
Converting admission test scores to IQ figures introduces errors on top of everything already discussed, and each one pushes in a direction that makes the resulting numbers look more authoritative than they are.
The tests measure overlapping but distinct constructs. Admission tests are built to predict academic performance and include substantial acquired knowledge, particularly vocabulary and mathematical content taught in school. Cognitive batteries include measures with no curricular content at all, such as matrix reasoning and processing speed. The correlation between them is substantial and far from unity, so a conversion assumes an equivalence the instruments do not have.
The norm groups are incompatible. An IQ score is a comparison against a general population sample. An admission test score is a comparison against other test takers, who are the highly selected group described in section 5. Converting a within-test-takers standing into a general population figure requires knowing where the test-taker population sits within the general one, which requires an assumption that is rarely stated and never justified.
Score scales change. Admission tests are periodically restructured and rescaled, and a conversion derived under one scale does not apply under another. Tables built at different times using different scale versions produce different IQ figures for the same field, and the discrepancy is a scaling artefact rather than a finding.
Finally, the conversion collapses the verbal and quantitative split into a single number, discarding the most informative feature of the data as covered in section 2.
The result is a table of precise-looking figures whose precision is entirely manufactured. The ordering it presents is real and available directly from the published score data, without the added layer that changes what it appears to be saying.
7 What This Means for Choosing a Major
People arrive at this data for a practical reason, usually either deciding what to study or evaluating a decision already made. Both deserve a direct answer.
The ranking is not a difficulty ranking. It reflects who applies to graduate programmes in each field, filtered several times. A field can be intellectually demanding while sitting mid-table because most of its graduates enter employment rather than continuing to doctoral study.
Your profile matters more than the ordering. Somebody with strong verbal and moderate quantitative ability will find a verbally demanding field easier going than a quantitatively demanding one, regardless of where either sits in a combined ranking. The domain profile is the relevant information and a single composite conceals it, which is why an assessment reporting domains separately is more useful for this question than one reporting a total.
The overlap is enormous. Field distributions overlap so heavily that individual position within a field matters far more than which field. A strong student in a mid-table field outperforms a median student in a top-ranked one comfortably, and this is the ordinary case rather than an exception.
Interest predicts outcomes that ability alone does not. Completion, performance, and persistence all depend on sustained engagement, and engagement is not distributed by ability. Choosing a field because it ranks highly, against your own interest, optimises a statistic at the expense of the outcome the statistic was supposed to indicate.
Nothing here is a ceiling. A field mean says nothing about whether a specific person can succeed in it. The distributions are wide, the filtering happens at multiple stages for reasons including preparation and opportunity, and treating a group average as a personal threshold inverts what the statistic means.
8 The Relationship to Earnings Is Weaker Than Assumed
These rankings are frequently repurposed as career guidance, and the mapping between them and financial outcomes is looser than the pairing implies.
Fields near the top of the test score ranking include several with modest earnings distributions. Philosophy, mathematics, and physics all place highly on the tests while producing graduates whose median earnings are well below those of fields lower in the ranking, particularly applied professional and business fields whose test performance is unremarkable.
The reason is that earnings are determined largely by what a field's labour market demands, and that demand tracks the economic value of specific skills rather than the general cognitive profile of the people who hold them. A field can select strongly for ability and lead into work that is poorly paid, and another can select weakly and lead into work that is highly paid, and both are common.
Where the two do connect is through subsequent education. Fields with high test performance frequently serve as routes into professional and doctoral programmes whose earnings outcomes are strong, which means the eventual financial result depends on the second step rather than the first.
The relationship between cognitive ability and income at the individual level is genuinely positive and considerably weaker than intuition suggests. It leaves most of the variation in earnings unexplained, and factors including field, sector, geography, negotiation, and opportunity account for far more.
Anybody using this data for career planning should therefore treat it as describing the composition of fields rather than their outcomes, and consult actual earnings data for the outcome question. The two answer different things and the ranking is not a proxy for the money.
9 The Confound Nobody Adjusts For
One factor distorts these rankings more than any other and is almost never mentioned: the proportion of test takers whose first language is not English.
Graduate programmes in engineering, computer science, mathematics, and the physical sciences enrol very large numbers of international students, and those students sit the same admission test as everybody else. Programmes in English literature, history, philosophy, and education enrol far fewer, because the disciplinary content is language-bound in a way that mathematics is not.
The verbal section is an English vocabulary and reading comprehension test. Somebody with excellent verbal reasoning operating in a second language scores below their ability by an amount that no adjustment recovers, for the reasons set out in Average IQ by Language. When a field's test-taker population includes a large share of such candidates, its verbal mean falls for reasons that have nothing to do with the reasoning ability of anybody in it.
This produces a specific and predictable distortion. Quantitative-heavy fields are penalised on the verbal section relative to their actual verbal ability, and humanities fields are advantaged relative to theirs, purely through language composition. The gap between physics and philosophy on verbal performance is therefore inflated by a factor that has nothing to do with either discipline.
The quantitative section is affected far less, because mathematical notation is close to language-independent and the reading load in quantitative items is light. This asymmetry is itself evidence for the mechanism: if the difference were about ability rather than language, it would appear on both sections rather than concentrating on the one that requires English.
Nobody adjusts for this, because doing so would require knowing each test taker's language background, which the published summaries do not report. The effect is therefore baked into every ranking derived from the data, and it operates in a consistent direction that anybody reading such a ranking should hold in mind.
10 Why the Ordering Barely Moves
The stability of this ranking across decades is frequently cited as evidence that it captures something fundamental. It is better understood as evidence that the filters have not changed.
The sorting mechanisms are institutional and slow. Prerequisite structures, which subjects require which earlier courses, have been broadly constant. The reputation of fields as demanding or otherwise shapes who considers them and persists across generations. The proportion of each field's graduates continuing to postgraduate study is determined by labour market structure, which changes slowly.
Where the ordering has moved, it has moved for identifiable structural reasons rather than because the fields changed intellectually. Computer science rose as it grew from a small specialism into a large discipline, which changed who entered it. Fields that expanded rapidly saw their means drift toward the overall average, which is the mechanical consequence of admitting a larger and therefore less selected population. Fields that contracted moved the other way for the same reason.
That last pattern is the clearest evidence that the ranking tracks selection rather than substance. A field does not become intellectually easier by enrolling more students, and its mean falls anyway. A field does not become harder by shrinking, and its mean rises.
The implication for reading a current ranking is that it describes the present configuration of a sorting system, not a permanent property of the subjects. Anybody who studied a field a generation ago belongs to a differently selected population than somebody studying it now, and comparing them through their field's current position is comparing two different things.
11 Using This Data Responsibly
There are legitimate uses for field-level score data, and they are narrower than the uses it is commonly put to. Separating them is worth doing explicitly.
Legitimate. Understanding the composition of applicant pools, which is directly useful to admissions offices comparing candidates across fields. Examining how a field's intake has changed over time, which is informative about growth and selection. Setting expectations about the preparation an incoming cohort brings.
Not legitimate. Inferring anything about a specific person from their field, which the overlap in section 2 rules out. Ranking fields by intellectual worth, which the selection problem in section 4 rules out. Screening candidates by degree subject, which compounds the language confound in section 9 into something with a discriminatory effect regardless of intent.
That last point deserves emphasis for anyone in a hiring position. Using field of study as a proxy for cognitive ability is both statistically weak, because within-field variation dwarfs between-field variation, and systematically biased, because the ranking is partly an artefact of who sits an English-language test. A selection process built on it would perform worse than one ignoring the field entirely, while carrying legal exposure that a poorly performing process does not usually justify.
The general principle is the one that runs through every group-average article: a mean describes a population and predicts an individual poorly. Where a decision concerns a person, the relevant evidence is about that person, and where none exists the correct response is to gather some rather than to substitute a group statistic for it.
12 A More Useful Version of the Question
The question people are really asking is usually not about fields. It is about themselves, and there is a better way to answer it.
Whether a demanding field suits you is a question about your profile, and specifically about which domains are strong relative to the others. Somebody whose fluid reasoning is well ahead of their crystallised knowledge is positioned differently from somebody with the reverse pattern, and both are positioned differently again from somebody with a large quantitative advantage.
This is exactly the information that a domain-level assessment provides and that no field ranking contains at any level of detail. A composite figure cannot distinguish between profiles, and the field ranking is a composite of composites.
ACIS reports six domains separately, including quantitative reasoning as its own index rather than folded into others, each with a percentile and confidence interval against a stated reference group. For the specific question of which kind of academic work will feel effortful and which will not, that structure answers directly what a field average answers by implication.
The limitation should be stated as clearly as the use. Administration is unsupervised, conditions cannot be verified, and no institution is obliged to accept the result. It informs a personal decision, which is the decision at issue here, and it is not admissions evidence.
It is also worth being clear about what a profile does not settle. Somebody whose quantitative index sits below their verbal one is not thereby unsuited to a quantitative field, because preparation, interest, and the amount of work somebody is willing to put in all move outcomes substantially and none of them appears in a cognitive profile. What the profile tells you is where the effort will land, not whether the effort will succeed, and those are different questions that a field ranking conflates entirely.
13 FAQ: IQ and Academic Fields
Which major has the highest average IQ?
No major has a measured IQ average. On graduate admission test data, physics, mathematics, philosophy, and economics consistently produce the highest combined scores.
Why does philosophy rank so highly?
Its disciplinary content closely matches what verbal reasoning items measure, it attracts students choosing it without vocational incentive, and its graduate programmes filter unusually hard.
Where does the data come from?
Primarily from the Educational Testing Service's annual summary of GRE takers by intended graduate field, which is public and methodologically documented.
Is the GRE an IQ test?
No. It is an admission test built to predict academic performance, containing substantial acquired knowledge. It correlates with general ability without being interchangeable with it.
Do demanding majors make you smarter?
Education raises cognitive test performance measurably, roughly one to five points per year of schooling. That is an effect of schooling in general and is far smaller than the between-field differences.
So the ranking is mostly selection?
Yes. By the time somebody appears in graduate admission data they have passed through four or five successive filters, each correlated with the abilities the test measures.
What is range restriction and why does it matter?
The population sitting these tests is heavily selected, so its variance is compressed. Field means within it cannot be extrapolated to everybody who studied those fields.
Why does uneven restriction distort comparisons?
Fields where most graduates continue to postgraduate study contribute their full distribution. Fields where few continue contribute only their upper tail, which makes them look stronger.
Why do different tables give different IQ numbers?
Because admission tests are periodically rescaled and each table applies its own undisclosed conversion. The discrepancies are scaling artefacts rather than findings.
What gets lost in converting to a single IQ figure?
The verbal and quantitative split, which is the most informative feature of the data. Two fields can reach the same total through entirely different profiles.
Is this a ranking of how hard majors are?
No. A field can be intellectually demanding while sitting mid-table because most of its graduates enter employment rather than continuing to doctoral study.
Should I pick my major based on this?
No. Your own profile matters far more than the ordering, and interest predicts completion and performance in ways ability alone does not.
How much do field distributions overlap?
Enormously. A strong student in a mid-table field routinely outperforms a median student in a top-ranked one, and that is the ordinary case.
Does a low field average mean I cannot succeed there?
No. A group average is not a personal threshold, and treating it as one inverts what the statistic means.
Do the top-ranked majors earn the most?
No. Several top-ranked fields produce median earnings well below fields whose test performance is unremarkable, because earnings track labour market demand rather than cognitive profile.
How strong is the link between ability and income?
Genuinely positive and considerably weaker than intuition suggests, leaving most variation in earnings unexplained by ability.
What about SAT data by intended major?
It describes school leavers stating a preference rather than anybody who completed a degree, so it is less filtered and produces a compressed version of the same ordering.
Does studying a subject improve the matching test section?
Yes, field-specific training produces field-specific improvement. Whether that constitutes general ability gain or a well-practised specific skill is the transfer question.
Are the differences between adjacent fields meaningful?
Frequently not. They are often small relative to within-field variation, while a ranked list presents them as though the ordering were sharp.
What should I look at instead?
Your own domain profile, showing where your abilities diverge, which no field ranking contains at any level of detail.
Is the published score data better than the IQ tables?
Yes. It carries the same ordering without a conversion that manufactures precision and discards the verbal and quantitative distinction.
14 Best Next Step
The ordering by field is real, stable, and mostly about who applies. It is not a difficulty ranking, not an earnings guide, and not a statement about what any individual can do.
For the domain structure that makes a profile more informative than a composite, read Cognitive Domains. For why group averages say little about individuals, read Average IQ by State. For what training does and does not change, read Can You Improve Your IQ?. For your own profile across six domains, take the assessment.
Field-level score data comes from the testing organisation that publishes it. Claims about selection, range restriction, and education effects come from the methodological and meta-analytic literature.
Educational Testing Service. GRE General Test score statistics, including the annual snapshot of test takers by intended graduate major. The primary data behind every published major ranking.
Educational Testing Service. Interpreting GRE scores. What the score scales represent, the reference population, and the limits on comparisons across score scale revisions.
Ritchie, S.J. & Tucker-Drob, E.M. (2018). How much does education improve intelligence? A meta-analysis. Psychological Science, 29(8), 1358-1369. Quasi-experimental estimate of roughly one to five IQ points per additional year of schooling.
Kuncel, N.R. & Hezlett, S.A. (2007). Standardized tests predict graduate students' success. Science, 315(5815), 1080-1081. What admission tests predict and the role of range restriction in estimating those relationships.
Kuncel, N.R., Hezlett, S.A. & Ones, D.S. (2001). A comprehensive meta-analysis of the predictive validity of the Graduate Record Examinations. Psychological Bulletin, 127(1), 162-181. Establishes the relationship between GRE performance and both academic outcomes and general cognitive ability.
McGrew, K.S. (2009). CHC theory and the human cognitive abilities project. Intelligence, 37(1), 1-10. The distinction between crystallised knowledge and fluid reasoning, which is why admission tests and cognitive batteries are not interchangeable.
National Center for Education Statistics. Digest of Education Statistics. Degree completions and graduate enrolment by field, which determine what fraction of each field's graduates appear in admission test data.
United States Bureau of Labor Statistics. Occupational Outlook Handbook. Earnings data by occupation, which is the correct source for the career question these rankings are frequently misused to answer.
Voncken, L., Albers, C.J. & Timmerman, M.E. (2019). Improving confidence intervals for normed test scores. Behavior Research Methods. Open access. Why differences between adjacent field means are frequently smaller than the uncertainty around them.
Crawford, J.R., Garthwaite, P.H. & Slick, D.J. (2009). On percentile norms in neuropsychology: proposed reporting standards. The Clinical Neuropsychologist, 23(7), 1173-1195. Why converting a standing within one reference population into another requires assumptions that are rarely stated.
Buros Center for Testing. Mental Measurements Yearbook. Independent evaluation of what admission tests and cognitive batteries respectively measure and support.
Take the assessment
You get a profile, not a number
ACIS measures six CHC domains across 20 subtests and reports each one with its own normed score and confidence interval, so you can see where you are strong and where you are not.