Institution Level

Average IQ by university

The figure most people carry in their heads for university students, somewhere between 115 and 130, was measured when almost nobody went to university. A meta-analysis of 106 Wechsler samples tested between 1939 and 2022 puts the current average at about 102. Institutions still differ from each other, and this page shows how much, in bands rather than in named numbers.

Six graduates in black caps and gowns with red tassels standing in a row outdoors, smiling and holding rolled diplomas tied with red ribbon.
In 1940, 4.6 percent of Americans aged 25 and over had completed four or more years of college. By 2021 the figure was 37.9 percent.

0 Quick Answer

There is no defensible single IQ figure for a named university, and the average across all undergraduates is now close to the population mean rather than a standard deviation above it. Bob Uttl, Victoria Violo, and Lacey Gibson meta-analyzed 106 samples of college and university students, 9,902 students in total, all tested on a Wechsler adult scale between 1939 and 2022. Their headline result is that the average undergraduate today scores about 102, and that undergraduate scores fell by roughly 0.2 points per year across that period. The full preprint of that meta-analysis carries the sample by sample table.

The mechanism is not mysterious and it is not a claim about anyone getting worse. In 1940, according to the US Census series the authors used, 4.6 percent of Americans aged 25 and over had completed four or more years of college. By 2021 that figure was 37.9 percent. When a selection threshold admits nine times as large a share of the population, the group behind the threshold moves toward the population it was drawn from. The old 115 figure was a real measurement of a much smaller and much more selected group.

Institutions still differ, and the difference is large. Uttl and colleagues report that the spread in admitted student test scores between the least and most selective American institutions exceeds three standard deviations. What that spread does not support is a table of precise numbers attached to named campuses, because every circulating table of that kind is back converted from median admission test scores, and the page on why score conversions mislead sets out in detail why that conversion fails. This page gives selectivity bands instead, with the derivation stated and a caveat inside every row.

102

Mean measured score for undergraduates today across 106 Wechsler samples, Uttl, Violo, and Gibson, 2024.

0.2 per year

Rate of decline in undergraduate sample means relative to the general population from 1939 to 2022.

.37

Correlation between an institution's estimated 2021 admitted student SAT and the measured ability of student samples drawn from it, across 80 of the 106 samples.

1 The Finding That Changed the Answer

The number that anchored this question for eighty years came from data collected when university education was a privilege of a small minority, and the people repeating it were citing sources that cited sources. Uttl, Violo, and Gibson traced the chain in their introduction, and the trace is the most useful part of the paper. Kaufman and Lichtenberger's 2005 text gives a college graduate mean of 115 and cites Matarazzo 1972, Jensen 1980, and Reynolds and colleagues 1987. Matarazzo's own table describes itself as based on clinical experience and offered as a working rule of thumb. Jensen's figures were compiled by Cronbach in 1960, who in turn drew on sources published between 1930 and 1958.

The meta-analysis then does the thing nobody had done: it collects the actual administrations. The inclusion criteria required samples of ordinary undergraduates tested with a Wechsler adult scale, and the search yielded 106 samples covering 9,902 students. The composition by instrument is 18 Wechsler-Bellevue samples, 28 WAIS samples, 40 WAIS-R samples, 17 WAIS-III samples, and three WAIS-IV samples. Full Scale IQ was reported directly for 100 of the 106 and imputed by regression for the remaining six.

The meta-regression on year of testing, fitted with a random effects restricted maximum likelihood estimator, gives a slope of negative .173 points per year for the 102 United States samples, with an R squared of .216 and a moderator test of QM(1) equal to 27.103, p below .0001. After adjusting each sample for the Flynn effect at the conventional rate, the slope steepens to negative .192 per year with an R squared of .242. Both readings land on the same conclusion, and the paper's own summary is that undergraduate students today average about 102.

The publication history matters and the authors document it themselves in a notice at the front of the preprint. Frontiers in Psychology accepted the paper on 4 January 2024, published the abstract, produced proofs, and then on 6 February 2024 emailed the authors rejecting the manuscript, citing overstated claims reported to its research integrity team. The authors dispute the rejection in detail. The peer reviewed version now sits at ScienceOpen under DOI 10.14293/S2199-1006.1.SOR.2024.0002.v1. Anyone citing this work should know that history, and should read the paper rather than the headline about it.

2 Why the Old Number Was Right Once

Nothing about the 115 figure was fraudulent when it was produced, and understanding why it stopped being true is the whole content of this page. Two forces moved it, and they are independent of each other. The first is the expansion of participation. The second is the Flynn effect, which shifts what any given raw performance is worth against a contemporary reference group.

Take participation first. The Census Bureau historical time series on educational attainment gives the shares Uttl and colleagues used. Among Americans aged 25 and over, high school completion rose from 24.1 percent in 1940 to 91.1 percent in 2021. The share with one to three years of college rose from 10.0 to 63.2 percent. The share with four or more years of college rose from 4.6 to 37.9 percent. Graduating from university today is more common than finishing high school was in 1940.

The arithmetic consequence is unavoidable, and the paper states it bluntly. If 80 percent of a population entered undergraduate education and that 80 percent averaged 115, the remaining fifth would have to average 40 for the whole population to average 100. Since the population average is fixed at 100 by the design of the scale, expanding the selected group has to pull its mean down toward the middle. This is a property of how scores are normed against a reference population, not a claim about declining ability in any absolute sense. The step that converts raw performance into a standard score, and the reason that step forces the population mean to stay at 100, is worked through on the page on how an IQ is calculated.

The second force is the drift in norms. Test scores rose at roughly three points per decade across the twentieth century, so a score of 115 obtained on a scale normed in 1950 is worth substantially less against a scale normed in 2022. Uttl and colleagues work the example through the Wechsler series: across the 53 years between the original WAIS and the WAIS-IV, Full Scale IQ rose by 13.3 points, or 0.25 per year, in samples given both. The Flynn effect page covers the phenomenon and the arguments about what drives it.

The normative samples themselves show the drop directly. From the WAIS-R norming, averaging 1978, to the WAIS-IV norming in 2007, the mean for adults with 16 or more years of education fell from 115.3 to 107.4, which is 0.27 points per year. For adults with one to three years of college, it fell from 107.4 to 101.4. Two independent lines of evidence, one from published normative samples and one from a meta-analysis of student administrations, arrive at the same place.

3 Why This Page Has No Table of Named Universities

Every circulating table of average IQ by named university is built the same way, and the construction is the part that fails. Somebody takes the published median SAT or ACT of admitted students at an institution, runs it through a conversion equation or a percentile match, and prints the result as that institution's average IQ. The output looks like measurement. It is a transformation of an admissions statistic.

This site already argues at length that admission test conversions mislead, and it would be incoherent to publish a table that depends on them. The reasons are set out on the page on why score conversions mislead, and the shortest version is that percentile matching assumes two reference groups are comparable when they are not, regression equations shrink toward the mean of the study sample they were fitted in, and reverse conversions are not algebraically valid because a regression predicting IQ from SAT is not the inverse of a regression predicting SAT from IQ. The same argument in the field of study context appears on the college major page.

There is a second problem specific to institutions, and it has gotten worse. A published median admission test score now describes only the subset of admitted students who chose to submit one. Uttl and colleagues flag this directly as a limitation of their own selectivity analysis, noting that the SAT figures they used reflect not all admitted students but only those who elected to report scores. As more institutions moved to test optional admission, the reporting subset became more selected than the class it is supposed to describe, which biases the median upward by an unknown amount that varies by campus.

Third, the institutional figures are medians of admitted students, and the people who enroll are not the people who were admitted. Highly selective institutions admit large numbers of students who enroll somewhere else. The published median therefore describes an offer pool, not a student body, and the gap between the two differs by institution in ways nobody publishes.

What this page will not doNo figure on this page is attached to a named university, and none is derived by converting an admission test score into an IQ. The bands below come from measured administrations of Wechsler scales to student samples, grouped by the selectivity of the institutions those samples were drawn from. Where the band edges are inferred rather than measured, the section that follows says so and shows the arithmetic.

4 The Selectivity Band Table and How It Was Built

Institutions can be grouped by how much of the applicant pool they admit, and the measured evidence supports a wide overlapping band for each group rather than a point. The bands below are defined by admission rate and by the admitted student test score midpoint that the National Center for Education Statistics collects in the Integrated Postsecondary Education Data System, which covers roughly 2,000 American institutions. The ability column is a band, not an estimate of any campus, and every row carries the reason it cannot be tightened.

Selectivity bandWhat defines it in the IPEDS dataBand the evidence supports for a student group meanHow the band was derivedCaveat carried by this row
Open admissionAdmits essentially all applicants who meet a minimum requirement; frequently reports no admission test scores at allAt or slightly below 100Uttl and colleagues report that a large proportion of IPEDS institutions have admitted student score midpoints below the College Board nationally representative mean, which corresponds to a group mean at or under the population averageThese institutions are the least represented in the meta-analysis, so the band rests on the admission test distribution rather than on many measured administrations
Broadly accessibleAdmission rate above roughly 75 percent, admitted student SAT midpoint near the national averageRoughly 98 to 105Anchored on the meta-analytic mean of about 102 and on the WAIS-IV normative mean of 101.4 for adults with one to three years of collegeThe WAIS-IV anchor is a 2007 norming of all adults at that attainment level, including people who attended decades earlier, so it is likely to overstate today's enrolled students
SelectiveAdmission rate roughly 25 to 75 percentRoughly 101 to 110Interpolated between the meta-analytic mean and the upper end of the observed sample means, using the .37 correlation between institutional selectivity and measured sample abilityAn r of .37 leaves most of the variance in sample means unexplained by selectivity, so two institutions in this band can differ in either direction
Highly selectiveAdmission rate roughly 10 to 25 percent, admitted student SAT midpoint well above the national averageRoughly 106 to 116Upper portion of the observed range of sample means in the meta-analysis, which the authors describe as running from below 100 to over 120Few of the 106 samples came from institutions in this band, and those that did were tested in different decades against different norms
Most selectiveAdmission rate under 10 percent; the highest admitted student SAT midpoint in the 2020 to 2021 IPEDS file is 1555 at the California Institute of Technology, with a 6.7 percent admission rateRoughly 110 to 122 and aboveTop of the observed range of student sample means, cross checked against the three standard deviation spread in admitted student scores that Uttl and colleagues measured across institutionsThis is the band with the fewest measured administrations and the strongest reporting bias, since score submission at these institutions is heavily self-selected

Read the fourth column before the third. The bands are wide because the evidence underneath them is wide, and the overlap between adjacent rows is the honest part of the table. A student group at a highly selective institution and a student group at a selective one can produce the same mean, and in the meta-analytic data they sometimes did.

A table is only as good as the arithmetic behind it, so here is the arithmetic, including the steps that are inference rather than measurement. Three published quantities do the work, and one derivation combines them.

  • The measured center. Across 106 samples the meta-analytic estimate for undergraduates today is about 102, with sample means running from below 100 to over 120 in the authors' own description of the distribution.
  • The spread in institutional selectivity. Using the 2020 to 2021 IPEDS admissions file, Uttl and colleagues report that the range between the least and most selective institutions exceeds three standard deviations on the admission test scale, which they translate as the equivalent of 45 points on the IQ scale. That translation is theirs, not this page's.
  • The link between the two. Across the 80 samples for which institutional score data were available, the correlation between an institution's estimated 2021 admitted student SAT and the Flynn adjusted measured ability of student samples from that institution was .37, p below .001. Adding institutional SAT as a second moderator alongside year of testing raised the model R squared from .242 to .325, meaning selectivity explained an additional 6 percent of the variance.

The derivation follows from the third quantity. A correlation of .37 means that a three standard deviation difference in institutional selectivity predicts an expected difference of about 1.1 standard deviations in student group means, which on a 15 point scale is around 17 points. That is a derived figure, calculated here from the published correlation and the published spread, not a number any of these authors printed. It is consistent with the observed range of sample means, which runs about 20 points wide, and it is what sets the width of the bands in the table above.

The important property of a correlation that size is that it constrains group means and says almost nothing about individuals. Selectivity accounts for roughly 14 percent of the variance in these sample means. The rest sits in the composition of the particular class, the field mix at that institution, the year of testing, and ordinary sampling noise across samples that are mostly small.

What is not known hereNobody has administered a current full length adult battery to representative samples of enrolled students at a set of named institutions in the same year under the same conditions. Until that exists, every institution level figure is either an inference from admission tests or a meta-analytic estimate assembled from administrations decades apart. This page is the second kind, and the bands are wide because that is what the second kind supports.

5 What the SAT and g Correlation Actually Supports

The evidence that admission tests measure general ability is strong, and it is exactly that evidence which forbids the point estimates people build from it. The foundational study is Meredith Frey and Douglas Detterman's Scholastic Assessment or g?, published in Psychological Science in 2004, volume 15, pages 373 to 378.

In Study 1 they took the National Longitudinal Survey of Youth 1979, extracted a general factor from the ten subtests of the Armed Services Vocational Aptitude Battery across 11,878 respondents, converted it to an IQ metric, and correlated it with SAT scores for the 917 people who had both. The correlation was .820. Adding a squared SAT term to correct for nonlinearity raised the multiple R to .857. The standard error of prediction for their equation was 5.94 points, which they contrasted favorably with the 11.4 point standard error of the demographic method clinicians used to estimate premorbid ability.

Study 2 is the one that matters for this page. They gave Raven's Advanced Progressive Matrices to undergraduates at a private university and obtained SAT scores from admissions records. Among the 103 usable cases the correlation was .483, corrected to .72 for restriction of range, and the standard error of prediction for the equation they fitted there was 9.76 points. They were explicit that Equation 1 could not carry over to recentered scores and that Equation 2 was most useful at the high end of the distribution and had to be used with caution.

The corrections came quickly and continued. Brent Bridgeman of Educational Testing Service published a comment titled Unbelievable Results When Predicting IQ From SAT Scores in Psychological Science in 2005, volume 16, pages 745 to 746, with a reply from Frey and Detterman at page 747; the exchange sits behind a paywall and this page does not summarize its content. Beaujean and colleagues published a validation of the two equations against the Reynolds Intellectual Assessment Scales in Personality and Individual Differences in 2006, volume 41, pages 353 to 357. Koenig, Frey, and Detterman extended the work to the ACT in Intelligence in 2008, volume 36, pages 153 to 160.

Frey's own fifteen year retrospective is the cleanest statement of where the literature landed. Writing in the Journal of Intelligence in 2019, she reports that researchers have confirmed the principal finding, yielding correlations between intelligence and the SAT of roughly 0.5 to 0.9 depending on the sample and on how intelligence is defined. A relationship whose published estimate spans four tenths of a correlation coefficient supports the statement that these tests overlap heavily with general ability. It does not support a single conversion constant, and a table that prints one is hiding the range inside a decimal point.

6 Restriction of Range Inside a Selective Institution

The single most useful demonstration of what selection does to a correlation is sitting inside the Frey and Detterman paper, and it is usually skipped. Their two studies used the same predictor and produced correlations of .820 and .483. The difference is not measurement error. It is who was in the room.

The Study 1 sample had a mean SAT of 854 with a standard deviation of 226, against a standardization standard deviation of 200. The Study 2 sample, drawn from a private university, had a mean of 1372 with a standard deviation of 119. Compressing the spread of a predictor to roughly half its national value cut the observed correlation by more than 40 percent. The authors corrected the Study 2 figure back up to .72 to estimate what it would have been in a less restricted sample, and they were careful to describe that as an estimate for a broader college population rather than a property of their own group.

This is the mechanism that makes within institution comparisons treacherous. Inside a single selective university, the admissions process has already removed most of the variance in the very quantity people want to correlate with something. Any relationship measured there will be smaller than the same relationship in the population, and the shrinkage is a function of how much the range was cut, not of how real the relationship is. The general machinery, and what a correlation does and does not license, is covered on the reliability and validity page.

There is a second consequence people miss. Restriction of range also flattens the ceiling. Frey and Detterman note that the Raven's scores in Study 2 showed a ceiling effect, which further suppressed the correlation. Measurement precision falls at the top of any scale because fewer items discriminate there and fewer people in the reference sample anchor the conversion, which is why the score to percentile relationship gets coarser as you move outward from the middle.

Applied to this page's question, the finding is this. Selectivity predicts group means at the level of institutions with an r of .37, but inside any one institution the same predictor loses most of its power, because the institution has already used it. That is why an institutional average, even a correct one, tells an individual student almost nothing about themselves.

7 The Spread Inside a Single University Is Larger Than People Expect

People imagine a selective university as a narrow slice near the top of the distribution, and the measured spread inside the college educated population is nearly as wide as the population itself. The clearest published figures come from the Wechsler normative samples, which record educational attainment for every examinee.

In the WAIS-IV normative sample, adults with 16 or more years of education had a mean Full Scale IQ of 107.4 with a standard deviation of 13.9, and the middle 95 percent of that group ran from 80 to 135. Adults with 13 to 15 years of education had a mean of 101.4 with a standard deviation of 13.1, and a middle 95 percent from 76 to 127. The population standard deviation is 15 by definition. Selecting on university completion cut the spread by about one point of standard deviation, not by half.

80 to 135

Middle 95 percent of Full Scale IQ among adults with 16 or more years of education in the WAIS-IV normative sample, mean 107.4, standard deviation 13.9.

13.9 vs 15

Standard deviation within the college graduate group against the population standard deviation, a reduction of about one point.

80 to 146 plus

Range of converted scores among college graduates in the Wonderlic 1992 data, as reported by Uttl and colleagues.

Two things follow. First, any institutional average conceals a distribution roughly as wide as the one it was drawn from, so knowing the average tells you very little about where a particular student sits. Second, group differences that look large as averages are small compared with individual differences inside each group. A ten point gap between two institutional means is real and is dwarfed by the 55 point interval that the middle 95 percent of either institution occupies.

The same logic applies to profiles rather than composites. Two students with the same overall score can differ sharply across the cognitive domains, one carrying a verbal comprehension peak and a processing speed trough, the other flat. The composite describes overall standing accurately in both cases while the index pattern supplies the direction that the single figure cannot, which is the argument developed on the fluid versus crystallized comparison.

8 What Admissions Selects for Beyond Tested Ability

An admissions decision is a bet on completion and on institutional interest, and tested ability is one input among several that are not interchangeable with it. This matters here because it explains why institutional selectivity and measured ability correlate at .37 rather than at .8.

Start with what the test itself contributes. In the most recent College Board validity sample, covering nearly a quarter million students, adding SAT scores to high school grade point average increased predictive power for first year college grades by roughly 15 percent over grades alone, and improved the prediction of retention into the second year. Frey reports both figures in her 2019 review. That is a real contribution and a bounded one: most of the prediction comes from the high school record, which reflects four years of accumulated work habits as much as ability.

Then there is everything the test does not measure. Conscientiousness, study habits, and attitudes predict academic achievement consistently, as Kuncel and Hezlett documented and Frey summarizes. Grit, once treated as a major addition, turns out on meta-analytic evidence to predict achievement only modestly and not incrementally over cognitive ability and conscientiousness. None of these traits is captured by an admission test, and all of them are read, imperfectly, from the rest of an application.

Selection also runs on things that are not about the applicant at all. Frey cites Klugman's analysis of a nationally representative sample showing that high school resources, both programmatic ones such as advanced placement offerings and social ones such as the socioeconomic composition of the student body, shape which institutions students apply to in the first place. Selection cannot operate on people who did not apply.

Finally, institutional incentives feed back into the numbers. Frey notes that the correlation between average admission test scores and college ranking in one widely read American ranking is very nearly .9. When a published median is itself a ranked quantity, institutions have a direct reason to manage which scores get reported, which is the same reporting bias described earlier from a different direction. For the general relationship between measured ability and school performance, the academic achievement page carries the effect sizes.

9 Why International Comparisons Do Not Work Here

Almost everything on this page is American, and that is a limitation rather than an oversight. Of the 106 samples in the meta-analysis, 102 came from the United States and four from Canada. The authors state plainly that four Canadian samples were too few to examine the Canadian trend, and their selectivity analysis rests entirely on an American federal data system and an American admission test.

Three obstacles block a clean international table. The first is that the word university does not describe the same institution across systems. Some countries route a large share of post-secondary students into vocational and technical institutions counted separately, others count them inside the university sector, and the resulting participation rates are not comparable without adjustment. A national figure for university students is partly a statement about how that country's statistical agency draws the boundary.

The second is that admission machinery differs. The United States runs a national commercial admission test whose results are published per institution. Many systems select on national school leaving examinations, or on subject specific entrance examinations, or on a lottery among qualified applicants. There is no common currency, so the selectivity variable that carries this page's table does not exist in most countries.

The third is norm provenance. A score means something only against a reference sample, and reference samples are national. Uttl and colleagues make this point about Canada specifically, citing Longman and colleagues' 2007 analysis showing that the association between WAIS-III Full Scale IQ and educational attainment was much smaller in the Canadian population than in the American one. Their own expectation is that Canadian undergraduate means, measured against Canadian norms, would sit around 100 or 101 rather than higher, and they report a 2022 Canadian undergraduate sample at 103 using American Shipley-2 norms from 2008. Applying one country's norms to another country's students shifts the answer before any comparison begins, which is the same issue described on the norming page.

The practical consequence for a reader outside the United States is that neither the selectivity bands above nor any admission test conversion transfers. What transfers is the mechanism: participation expanded almost everywhere across the same eighty years, and wherever it did, the selected group moved toward the population it was drawn from.

10 The Harvard Question and the Ivy League Question

These are the two most common searches on this topic and both have the same honest answer, which is that no such measurement exists. No institution in the Ivy League administers an individually normed intelligence battery to its students and publishes the mean. What circulates instead is a conversion of a published admission statistic, produced by someone outside the institution, and the previous sections explain why that conversion does not survive scrutiny.

What can be said is bounded and still useful. Institutions in the most selective band admit under 10 percent of applicants and report admitted student score midpoints at the top of the national distribution; the highest midpoint in the 2020 to 2021 IPEDS admissions file was 1555 at the California Institute of Technology, paired with a 6.7 percent admission rate. Student groups drawn from institutions in that band sit at the top of the observed range of measured sample means, which Uttl and colleagues describe as reaching over 120. The band in the table above runs from about 110 to 122 and above, and the width is not hedging. It is what an r of .37 between selectivity and measured ability permits.

Three further points keep the answer honest. The admitted pool is not the enrolled class. The reported score median covers only students who submitted scores. And an institutional mean, however accurately measured, would still sit inside a distribution roughly 55 points wide, so it would not describe any individual graduate of that institution.

This is also why a university credential works poorly as an ability proxy for a specific person, which is the pattern behind most celebrity intelligence claims. A degree from a selective institution gets converted into an implied score, and the implied score gets repeated as though it had been measured. The page on Natalie Portman traces one widely repeated case where a Harvard credential is doing exactly that work, and the myths page catalogues several other claims that survived by repetition rather than by evidence.

The claim this page does not makeNothing above is a statement about any individual student at any institution. Group statistics describe distributions and never diagnose people, and a difference of a few points between two group means is real at the population level and invisible at the level of one person. ACIS has never assessed any named institution's student body, and no retrospective assessment of one is possible.

11 Measuring the Person Instead of the Institution

The question behind the search is almost always personal, and an institutional average is the wrong instrument for a personal question. If someone wants to know where they stand, the answer requires an administration against an adult reference frame, not a lookup against a campus.

Two routes exist and they answer different questions. A proctored session with a licensed psychologist using a current battery produces a result that institutions accept, because identity was verified, conditions were controlled, and a qualified examiner observed the performance. That is what a clinical process, an accommodation request, or a court needs, and nothing administered over the internet substitutes for it. The routes and costs are laid out on the where to take a test page and the professional versus online comparison.

A self-administered online assessment answers a different question, and the differences between a normed battery and a quiz are set out on the free versus validated comparison. ACIS is the second kind. It runs 20 subtests across six CHC domains for adults aged 16 to 90, reports a Full Scale IQ with six primary indices, and is sold as a one time payment in three tiers: Quick at 15 dollars, Optimized at 30 dollars, and Full Scale at 50 dollars.

For a reader arriving from this page, the Quick tier is usually the right starting point. It runs six subtests, Similarities, Vocabulary, Matrix Reasoning, Figure Weights, Digit Span, and Alphanumeric Sequencing, in about 45 minutes, and returns verbal comprehension, fluid reasoning, and working memory alongside a partial profile. The published reliability figures for the composites those subtests feed are an omega of .9745 for the verbal comprehension index with a standard error of measurement of 2.40 points, .9727 for fluid reasoning with a standard error of 2.48, and .9247 for working memory with a standard error of 4.12. The derivations sit in the technical manual.

ACIS is self-administered, it is not a clinical instrument, and it is not appropriate for diagnosis, hiring decisions, accommodation requests, or high IQ society admission. Its adult reference frame is documented in the technical manual. What the session involves is described on the adult test page, and what a resulting figure means in percentile terms is on the tool that converts a score to a percentile.

12 What the Evidence Actually Supports

The defensible claim is narrow and it is worth stating without decoration. Across 106 Wechsler administrations to undergraduate samples between 1939 and 2022, the mean fell by roughly 0.2 points per year, and it now sits near 102. That decline tracks a nine fold expansion in the share of Americans completing four or more years of college, from 4.6 percent in 1940 to 37.9 percent in 2021, and it is what any selection threshold does when it stops being selective.

Institutions differ, and the difference is measurable at the group level. Selectivity, measured through admitted student test scores, correlates .37 with the measured ability of student samples drawn from those institutions and adds about 6 percent of explained variance to a model that already includes year of testing. That supports bands roughly 10 points wide with substantial overlap between adjacent bands. It does not support a number attached to a campus.

The two errors this page is written against are opposite and equally common. The first is treating a university credential as a proxy for a score, which the meta-analytic evidence has now made untenable. The second is concluding from that evidence that institutions are all the same, which the three standard deviation spread in admitted student scores contradicts. Both errors come from wanting a point estimate where the data supports a distribution.

What follows for a reader is simple. An institutional average cannot tell you about yourself, because the spread inside any institution is nearly as wide as the spread in the population. If the question is personal, it needs a personal measurement, taken against an adult reference frame and reported with its measurement error. If the question is about institutions, the honest answer is a band with its derivation attached, which is what the table above is.

13 Sources Behind This Page

Every figure above comes from one of the following, and each is linked so the arithmetic can be checked rather than trusted.

  • Uttl, B., Violo, V., and Gibson, L. (2024). Meta-analysis: On average, undergraduate students' intelligence is merely average. Peer reviewed version at ScienceOpen, DOI 10.14293/S2199-1006.1.SOR.2024.0002.v1. Full text preprint, which also carries the authors' notice about the Frontiers in Psychology acceptance and subsequent rejection. Source of the 106 samples, the 9,902 students, the mean of about 102, the negative .173 and negative .192 per year meta-regression slopes, the .37 selectivity correlation across 80 samples, the R squared of .325 for the two moderator model, the WAIS normative sample figures by education level, and the IPEDS selectivity statistics including the 1555 midpoint and 6.7 percent admission rate at the California Institute of Technology.
  • Frey, M. C., and Detterman, D. K. (2004). Scholastic Assessment or g? The Relationship Between the Scholastic Assessment Test and General Cognitive Ability. Psychological Science, 15(6), 373 to 378, DOI 10.1111/j.0956-7976.2004.00687.x. Full text. Source of the .820 correlation across 917 respondents, the .857 multiple R after the nonlinearity correction, the 5.94 point standard error of prediction, the .483 and .72 correlations in Study 2, the 9.76 point standard error there, and the sample means and standard deviations of 854 with 226 and 1372 with 119.
  • Bridgeman, B. (2005). Unbelievable results when predicting IQ from SAT scores. A comment on Frey and Detterman (2004). Psychological Science, 16(9), 745 to 746, with discussion at 747. PubMed record. Cited here only for the existence and location of the published dispute.
  • Frey, M. C. (2019). What We Know, Are Still Getting Wrong, and Have Yet to Learn about the Relationships among the SAT, Intelligence and Achievement. Journal of Intelligence, 7(4), 26. Open access full text. Source of the roughly 0.5 to 0.9 range of published SAT and intelligence correlations, the quarter million student validity sample, the 15 percent gain in predictive power over high school grades alone, the near .9 correlation between average admission test scores and one American college ranking, and the summaries of the Kuncel and Hezlett and Klugman findings.
  • Koenig, K. A., Frey, M. C., and Detterman, D. K. (2008). ACT and general cognitive ability. Intelligence, 36, 153 to 160. Cited for the extension of the 2004 work to the ACT.
  • Beaujean, A. A., Firmin, M. W., Knoop, A. J., Michonski, J. D., Berry, T. P., and Lowrie, R. E. (2006). Validation of the Frey and Detterman (2004) IQ prediction equations using the Reynolds Intellectual Assessment Scales. Personality and Individual Differences, 41(2), 353 to 357.
  • United States Census Bureau, historical time series on educational attainment. Source of the 1940 and 2021 attainment shares.
  • National Center for Education Statistics, Integrated Postsecondary Education Data System. Source of the institutional admission rates and admitted student score midpoints used to define the selectivity bands.
  • ACIS technical manual, for the reliability and standard error figures quoted for the Quick tier composites. The adult reference frame is documented in the technical manual.

The professional framework governing all of this is explicit. The Standards for Educational and Psychological Testing (2014), published jointly by the American Educational Research Association, the American Psychological Association, and the National Council on Measurement in Education, require that a score interpretation be supported by evidence for the specific use proposed, that classification statements be reported with their measurement error, and that the limits of the reference sample be disclosed. The APA standards on test use make the same demand of anyone who reports a group statistic: state what the number supports, state what it does not, and do not let a group average stand in for a measurement of a person.

14 Frequently Asked Questions

What is the average IQ of university students?

About 102 on current evidence. Uttl, Violo, and Gibson meta-analyzed 106 samples of undergraduates totaling 9,902 students, all tested on Wechsler adult scales between 1939 and 2022, and report that the contemporary average sits close to the population mean of 100.

Why do so many sources say 115 or higher?

Because they are citing measurements taken when university education was rare. The chain runs back through textbooks to figures compiled in the 1930s through the 1950s, one of which describes itself as a working rule of thumb based on clinical experience rather than a study.

Has the ability of students actually fallen?

Not in an absolute sense. The share of Americans aged 25 and over with four or more years of college rose from 4.6 percent in 1940 to 37.9 percent in 2021. A group that admits nine times as large a share of the population necessarily has a mean closer to that population.

What is the average IQ of Harvard students?

No such measurement is published, because no Ivy League institution administers an individually normed intelligence battery to its students and reports the mean. Every circulating figure is a conversion of a published admission test statistic by someone outside the institution.

Why will this page not give a number per university?

Because the per university numbers in circulation are back converted from median SAT or ACT scores, and that conversion does not survive scrutiny. Percentile matching assumes non comparable reference groups are comparable, and a regression predicting IQ from SAT is not the algebraic inverse of one predicting SAT from IQ.

How strongly does university selectivity predict student ability?

At r equal to .37 across the 80 samples in the meta-analysis with institutional score data. Adding selectivity to a model that already contained year of testing raised the explained variance from 24.2 to 32.5 percent, so selectivity contributed about 6 percentage points.

How wide is the difference between the least and most selective institutions?

Uttl and colleagues report that the spread in admitted student test scores across IPEDS institutions exceeds three standard deviations, which they translate as the equivalent of 45 IQ points. Because selectivity correlates .37 with measured ability, the expected difference in student group means is far smaller than that.

Does the SAT measure intelligence?

It correlates heavily with it. Frey and Detterman found .820 between SAT scores and a general factor extracted from the Armed Services Vocational Aptitude Battery among 917 respondents, and Frey's 2019 review reports that later work put the correlation anywhere from roughly 0.5 to 0.9 depending on sample and definition.

So why not just convert an SAT score to an IQ?

Because the conversion carries an error band that the printed number hides. Frey and Detterman's own equations had standard errors of prediction of 5.94 and 9.76 points, which means a 95 percent interval roughly 24 to 38 points wide around any single converted score.

What is restriction of range and why does it matter here?

It is what happens to a correlation when the spread of one variable is cut. Frey and Detterman's national sample had an SAT standard deviation of 226 and produced r equal to .820; their private university sample had a standard deviation of 119 and produced .483 on the same predictor.

How much variation is there inside one university?

More than people expect. In the WAIS-IV normative sample, adults with 16 or more years of education had a mean of 107.4 with a standard deviation of 13.9, and the middle 95 percent ran from 80 to 135, against a population standard deviation of 15.

Are published admission test medians reliable?

Less than they used to be. As institutions moved to test optional admission, published medians came to describe only the admitted students who chose to submit scores. Uttl and colleagues flag this explicitly as a limitation of their own selectivity analysis.

Is the admitted student median the same as the enrolled student median?

No. Selective institutions admit many students who enroll elsewhere, so the published median describes an offer pool rather than a class. The size of the gap differs by institution and is not published.

What else does an admissions process select for?

Conscientiousness, study habits, and attitudes predict academic achievement consistently and are not measured by an admission test. High school resources also shape which institutions students apply to at all, so selection cannot operate on people who never applied.

How much does the SAT add over high school grades?

About 15 percent more predictive power for first year college grades, in a College Board validity sample of nearly a quarter million students, and it also improves the prediction of retention into the second year. Most of the prediction still comes from the high school record.

Can these figures be applied outside the United States?

No. Of the 106 samples, 102 were American and four Canadian, and the selectivity analysis rests on an American federal data system and an American admission test. Norms are national, and the word university does not describe the same institution across systems.

What happened with the Frontiers publication of this meta-analysis?

Frontiers in Psychology accepted the paper on 4 January 2024 and published its abstract, then on 6 February 2024 emailed the authors rejecting it. The authors dispute the rejection in a notice at the front of the preprint, and the peer reviewed version is now hosted at ScienceOpen.

Does a degree from a selective university indicate a high score?

It shifts the odds and nothing more. The distribution inside any institution is roughly as wide as the population distribution, so a credential is a weak proxy for an individual result, which is exactly the reasoning error behind most celebrity intelligence claims.

Do these group differences say anything about an individual student?

No. Group statistics describe distributions, never people. A difference of a few points between two institutional means is genuine at the population level and completely invisible in any single case.

What would it take to answer this question properly?

A current full length adult battery administered to representative samples of enrolled students at named institutions, in the same year, under the same conditions, with the norms disclosed. Nothing like that exists, which is why every institutional figure is either an inference from admission tests or a meta-analytic estimate.

Has ACIS measured any university's students?

No. ACIS has never assessed any named institution's student body and makes no claim about one. It is a self-administered online assessment for adults aged 16 to 90, not a clinical instrument, and it is not appropriate for diagnosis, hiring, or admissions decisions.

Take the assessment

You get a profile, not a number

ACIS measures six CHC domains across 20 subtests and reports each one with its own normed score and confidence interval, so you can see where you are strong and where you are not.

Free trial, no card required. Full report from $15.