Every test manual in clinical use prints a table of classification labels at the top of its score report. None of them says genius. The word entered intelligence testing through a single suggested table in 1916, survived in popular use after its author abandoned it, and has never been a band on a modern scale.
The word genius entered IQ testing through one suggested classification table and left through the manuals that replaced it.
0 Quick Answer
No intelligence test currently in professional use has a classification band labeled genius, so there is no threshold to clear. The number people have in mind is 140, and it is not invented. It comes from one place: a suggested classification table on page 79 of Lewis Terman's The Measurement of Intelligence, the 1916 manual for the Stanford revision of the Binet scale, whose top row reads "Above 140: Near genius or genius."
The complication is that Terman did not mean the row the way it has been read for a century. In the same passage he warned that the boundaries between his groups were "absolutely arbitrary, a matter of definition only," and a few chapters later he wrote that intelligence tests had not been in use long enough to let anyone define genius in terms of IQ at all. He then spent thirty-five years running the study that was supposed to settle the question, and the results argued against his own label.
Modern publishers print something else. Pearson's WAIS-IV score report labels the top of its range Very Superior. The WAIS-5, published in 2024, replaced that vocabulary with directional descriptors. The ACIS classification table runs from Average through Superior to four gifted bands and never uses the word either. What follows is where 140 came from, why it persisted with no manual behind it, and what the longitudinal evidence actually shows about very high childhood scores.
Above 140
The only row in a published test manual that ever carried the word genius, in Terman's 1916 Stanford-Binet classification.
Zero
Times the word appears in the WAIS-IV and WAIS-5 sample score reports that Pearson publishes.
5 points
The childhood Stanford-Binet gap between the most and least successful men in Terman's own cohort, reported in 1947.
The threshold has a single documented origin, and it is a paragraph of advice rather than a finding. Terman published The Measurement of Intelligence in 1916 as a guide for examiners using his revision of the Binet-Simon scale. Chapter VI asks what the scores imply in ordinary language, and answers with a table he introduces cautiously: since terms like these "are convenient and will probably continue to be used, it is desirable to give them as much definiteness as possible."
The table he then offers has seven rows. Above 140 is "Near genius or genius." From 120 to 140 is "Very superior intelligence." From 110 to 120 is "Superior intelligence." From 90 to 110 is "Normal, or average, intelligence." The three rows below that use the clinical vocabulary of 1916 for below average performance, language no publisher would print now. Every row is a suggestion offered on the basis of the cases Terman had tested, not a value derived from a distribution.
He said so directly. Two conditions had to be borne in mind, he wrote: that the boundary lines between the groups are absolutely arbitrary, and that the individuals in any one group do not form a homogeneous type. Earlier in the same book he had made the same point about the shape of the distribution, arguing there is no line of demarcation between the extremes and the so-called normal child, and that classifying people into bands is like sorting them into abnormally tall, normally tall, and abnormally short.
When Terman reached the section on the top of his own scale, he was more explicit still. Intelligence tests, he wrote, had not been in use long enough to enable anyone to define genius in terms of IQ. The two highest records he could report from his own files were a boy of eight with an estimated 155 and a boy of seven with 160, and he noted that the 160 was the highest score in the Stanford University records at that point. The number 140 was a place to put a word, not a discovery about where a category begins. Anyone tracing the history of intelligence testing arrives at the same conclusion.
2 What a 1916 Score Actually Measured
A 140 in 1916 and a 140 today are not the same quantity, because the arithmetic underneath them is different. Terman's scale produced a ratio IQ: an examiner established the child's mental age from the highest set of items passed, divided it by chronological age, and multiplied by 100. A nine year old performing like an average twelve year old scored about 133. The number described a rate of development against an age standard, and it was built for children, since mental age stops being a coherent unit once growth flattens in adulthood.
Modern scales abandoned that arithmetic. A contemporary score is a deviation score: raw performance is ranked against the norm sample for the test taker's own age band, and the rank is expressed on a scale with a fixed mean of 100 and a fixed standard deviation of 15. That change matters most at the extremes, because a ratio can drift with age in ways a percentile cannot, and because two ratio scores of 140 obtained at different ages do not describe the same rarity. The difference between the two systems is set out in the mental age article and again in the explanation of the 15 point standard deviation.
The practical consequence is that no direct translation exists between Terman's row and any band on a current instrument. People quoting 140 as a genius threshold are quoting a cutoff on a scale that was retired, using a metric that was replaced, from a manual whose own author withdrew the label. The number kept traveling because it is short and memorable, not because it survived the transition.
There is a second reason the comparison fails. The 1916 sample was small and local. Terman standardized on roughly a thousand California children and never claimed a census representative frame. The reference groups behind modern instruments are built to match national population characteristics, which is what makes a percentile from one of them portable. How that machinery works, and what it costs when it is skipped, is covered in the article on norming.
3 Terman Dropped the Word Himself
The clearest evidence that genius was never a real classification is that its author stopped printing it after twenty-one years. In 1937 Terman published Measuring Intelligence with Maud Merrill, the manual for the new revised Stanford-Binet. Search the full text and the word genius appears exactly once in the entire book, inside a bibliographic footnote citing the title of his own research series, Genetic Studies of Genius. It is never used to label a score.
The Wechsler scales never adopted it at any point. Neither did the Raven's family, the Woodcock-Johnson batteries, or any of the school age instruments that followed. Whatever the reasons in each case, the outcome across nine decades of test publishing is uniform, and it is not a matter of interpretation: the tables are public and the word is absent from them.
Instrument and source
Year
Top classification label
What it says about genius
Stanford revision of the Binet scale, Terman, The Measurement of Intelligence
1916
Near genius or genius, above 140
The only manual row that ever used the word, offered as an arbitrary convenience
New revised Stanford-Binet, Terman and Merrill, Measuring Intelligence
1937
No genius label anywhere in the manual
The word survives once, in a footnote citing a book title
WAIS-IV Score Report, Pearson
2008
Very Superior
Word absent from the published sample report
WAIS-5 Score Report, Pearson
2024
Directional descriptors, top band Extremely High
Word absent from the published sample report
ACIS public classification table
Current
Profoundly Gifted
Four gifted bands above 130, no genius band
Read the table as a record of what publishers were willing to put their name to. A classification label is a communication device that a clinician reads aloud to a person who has just spent two hours being tested. The vocabulary has moved steadily toward statements about where a score sits and away from statements about what kind of person produced it, which is the whole reason the 1916 row now looks so strange.
4 What Test Publishers Print at the Top Today
Two Pearson documents, both freely downloadable, show the change in vocabulary directly. The WAIS-IV sample score report presents a Composite Score Summary with a column headed Qualitative Description. The sample examinee's Full Scale IQ of 139 is described as Very Superior, and so are the Verbal Comprehension index at 145 and the Working Memory index at 133. Perceptual Reasoning at 123 is Superior.
The WAIS-5 sample score report, dated 2024, keeps the same column heading and changes every word in it. The sample examinee's Verbal Comprehension index of 124 is now Very high. The Full Scale IQ of 111 is Above average. The superior and inferior vocabulary that Wechsler inherited from the 1930s is gone, replaced by labels that say only how far from the middle a score falls. According to the WAIS-5 Technical and Interpretive Manual (Pearson, 2024), the band at 130 and above is Extremely High.
What was verified hereThe Very Superior and Very high labels above were read directly from the two Pearson sample reports linked in this section, and neither report contains the word genius. The Extremely High label for 130 and above comes from the WAIS-5 manual, which is not a public document, so treat it as publisher reported rather than independently checked.
The rest of the field looks the same. The Stanford-Binet Fifth Edition uses gifted and advanced language at the top of its table rather than genius, and Raven's Progressive Matrices 2 reports percentile standing against its norm group. The instrument pages on the WAIS-5, the WAIS-IV versus WAIS-5 comparison and Raven's 2 set out what each battery covers and how their reporting conventions differ. Those conventions diverge further than the vocabulary does: Raven's 2 returns a raw score and a percentile rank rather than an index profile, and the KABC-II NU produces two different global composites for the same child depending on whether the examiner interprets under the Luria or the CHC model, differences carried row by row in the inventory of instruments in current use.
ACIS follows the same convention for the same reason. Its public table runs Average for 90 to 109, High Average for 110 to 119, Superior for 120 to 129, then Moderately Gifted, Highly Gifted, Exceptionally Gifted and Profoundly Gifted. The full set of bands with percentile context is on the score chart.
5 Why 140 Survives With No Manual Behind It
Nothing enforces the number, which is precisely why it never dies. A threshold with an institution behind it can be audited. If a body admits members at a stated percentile, you can check the criterion, check the approved test list, and check whether a given score qualifies. The 140 line has none of that. No publisher prints it, no society uses it as a gate, no clinical guideline references it, so no one is ever in a position to correct it. There is also a structural reason the correction never arrives. A threshold printed in a test manual can be revised when the manual is revised, and every examiner who buys the new edition gets the new table. A threshold living in general circulation has no editions. Terman's 1937 revision reached the people who purchased it and nobody else, so the 1916 row kept running in the popular literature for another ninety years with no mechanism capable of retiring it. The same asymmetry explains why the discarded vocabulary of the lower rows in that table vanished quickly while the top row did not. Nobody has an incentive to repeat the offensive labels, and almost everybody has an incentive to repeat the flattering one.
Compare the thresholds that do have institutions. Mensa's published membership criterion is a score at or above the 98th percentile on an approved supervised test, which on the standard 15 point scale is about 130 rather than 140. Other societies set higher bars and publish them, with the practical difficulties that creates for measurement at the top. Both are covered on the Mensa page and the high IQ society requirements page.
The second reason is that 140 is doing a different job in ordinary speech. When someone says a person has a genius IQ, they are usually making a claim about visible accomplishment, not about a test result they have seen. The number is borrowed to sound specific. This is the same mechanism that produces the estimated scores attached to historical figures, which almost never trace back to an administered test, a pattern documented on the highest IQ ever page. When the claim is about the speaker rather than about somebody else, the borrowing is no better founded, because self estimates pooled across 93 studies and 36,833 participants correlate about .30 with measured ability, loose enough that a guess on its own spans some 55 points, which is the arithmetic laid out in the evidence on self estimated intelligence.
The third reason is repetition. Newspapers, quiz sites and study guides copied the 1916 row long after Terman withdrew it, and each copy became a source for the next. A search for the phrase today returns pages that agree with each other and cite nothing, which is what a claim looks like when it has been transmitted rather than checked. The myths page catalogues several other claims that survived the same way, repeated until repetition looked like evidence.
6 The Study Terman Built Around His Own Threshold
Terman did not leave the label as a claim, and that is the most useful thing about him. In 1921 he began the longitudinal project published as Genetic Studies of Genius, designed to follow a large group of very high scoring California schoolchildren for the rest of their lives. If a threshold identified people who would go on to do extraordinary things, a cohort selected on that threshold would eventually show it.
Recruitment ran in stages. Teachers nominated their brightest pupils, nominees sat a group test, and the strongest performers were then given an individual Stanford-Binet by Terman's assistants. The final cohort was 1,528 children, 856 boys and 672 girls, scoring at roughly 135 and above, with 140 as the target threshold for the core sample. They became known as the Termites, and the study ran for the rest of the century, making it one of the longest longitudinal projects in the history of psychology.
The design has a limitation worth naming before the results, because it shapes everything that follows. Selecting on a score means the cohort has almost no variation in the very quantity under test. Everyone inside it is far above average, so any comparison made within the group is a comparison across a narrow slice at the top of the distribution, not across the population. Restriction of range does not invalidate the study, but it determines what the study can answer: it can say what differs among very high scorers, and it cannot say what differs between them and everyone else.
The second limitation is the sampling frame. Nomination by teachers filtered for pupils who were visible, cooperative and doing well in school before any test was administered. The cohort therefore over represents children whose ability showed up in ways a 1921 California classroom noticed and rewarded, which is a different population from all children who would have scored above 140 had they been tested. The Terman study also appears in the background of the Oppenheimer page, where the same design is examined from a different angle.
7 What the Cohort Actually Did With Their Lives
The Termites did well and did not do what the threshold was built to predict. As a group they finished more education than their contemporaries, entered professional and managerial occupations at higher rates, earned more, and reported better health and marital stability. Terman and Melita Oden documented all of it across the follow-up volumes. On any ordinary reading of outcomes, a cohort selected at the top of a childhood distribution came out ahead.
What the group did not produce was a single Nobel laureate. Two boys who were screened during the selection process and turned away did: William Shockley, who shared the 1956 Nobel Prize in Physics for the transistor, and Luis Alvarez, who won it in 1968 for work in particle physics. The screening records are documented, and the fact has been examined in the peer reviewed literature rather than passed around as an anecdote.
Read the miss carefullyWarne, Larsen and Clark reported in Intelligence in 2020 that they simulated Terman's selection procedure and found a 53 to 83 percent probability that it would catch neither future laureate, purely from base rates. Nobel Prizes are rare enough, and the threshold was high enough, that missing both was the expected outcome. The episode is evidence about the limits of predicting rare events, not evidence that the test was broken.
Both readings can be held at once, and the honest version keeps both. A high childhood score was genuinely informative about the ordinary run of adult outcomes, which is why the cohort averages came out where they did. It was close to useless for identifying the handful of people who would transform a field, because that outcome is too rare for any cutoff to select on. The floor of the distribution matters here as much as the ceiling. Oden observed in her 1968 report that most of the men classified as least successful still equalled or exceeded the vocational standing of unselected men of the same age. A cohort picked from the top of a childhood distribution produced very few genuinely poor adult outcomes, and that is the finding which survives every criticism of the study's design. What it did not produce was the tail the threshold was built to catch. The prediction held where outcomes are common and failed where they are rare, and those are not contradictory results. They are one result observed at two very different base rates. The general relationship between measured ability and life outcomes is covered on the success page and the income page, and the same argument appears with different sources on the Tyson page.
8 The Comparison Terman Ran Inside His Own Cohort
The strongest evidence about high scores comes from the one comparison Terman ran that nobody quotes. By 1940 the cohort was old enough to evaluate. Three judges, including Terman and Oden, examined the records of 730 men aged twenty-five and over and picked out the 150 most successful, called the A group, and the 150 least successful, called the C group, roughly the top and bottom fifth. They then compared the two on some 200 items of information gathered since 1921.
Chapter XXIII of The Gifted Child Grows Up, the 1947 volume, reports the test scores in Table 100. On the 1922 Stanford-Binet, the A group averaged 155.0 and the C group averaged 150.0, a gap of five points. On the Terman Group Test taken the same year, the two groups were 143.2 and 142.3, a difference the authors did not treat as reliable. Terman's own summary of the table is the sentence the popular literature skips: "Where all are so intelligent, it follows necessarily that differences in success must be due largely to nonintellectual factors."
Melita Oden repeated the exercise twenty years later and reported it in Genetic Psychology Monographs in 1968, reselecting 100 men at each end. Childhood Stanford-Binet means were 157.3 for the A group and 149.7 for the C group, a difference of 7.6 points that reached significance. Oden then listed the reasons not to lean on it: only about two thirds of each group had a Binet score, the absolute gap was small, and 17 percent of the C group scored above the A group mean. Her conclusion was blunt. It seemed to her highly unlikely that the large achievement differences between the groups could be attributed to the small intelligence score differences.
One measure did separate the two groups cleanly, and it is the one that cannot be read as a cause. The Concept Mastery Test, a demanding adult verbal battery built for people at the top of the range, was administered to the cohort in 1940, eighteen years after selection. The A group averaged 112.4 and the C group 94.1, a difference the authors treated as reliable. They then declined to lean on it, for two reasons stated in the same paragraph. The first is that both means were extraordinary in absolute terms: the C group score was equalled by only about 15 percent of students at superior universities, and the A group mean sat not far below that of doctoral candidates at leading institutions. The second is timing. By 1940 the A men had accumulated far more schooling than the C men, and Concept Mastery performance is influenced by education. Oden made the same point again in 1968, when Form T of the test produced 147 for the A group, 130 for the C group, and 119 for a comparison group of doctoral candidates at a leading university. The measure that best distinguished the two groups was taken after two decades of divergent education, which makes it as much a consequence of the outcome as a predictor of it.
155.0 vs 150.0
Mean childhood Stanford-Binet IQ for the most and least successful men, Terman and Oden 1947, Table 100.
17 percent
Share of the least successful group scoring above the most successful group's mean, Oden 1968.
92 vs 40 percent
College graduation rates for the two groups by 1960, from the same follow-up.
The size of the outcome gap is what makes the score gap look small. In Oden's 1960 data the median total family income was 33,125 dollars for the A group and 8,500 dollars for the C group. Educational attainment diverged just as sharply. Two groups separated by five to eight points of childhood score, both far above the 140 line, ended up four times apart in income. The pattern is consistent with the broader literature on measured ability and academic achievement, which finds real but partial prediction.
9 Cox 1926 and the Numbers Attached to Dead People
The other historical source people cite is Volume II of the same series, and its method deserves to be read before its results are. Catharine Cox published The Early Mental Traits of Three Hundred Geniuses in 1926 as her doctoral work under Terman. She took eminent historical figures from James McKeen Cattell's ranked list of a thousand names, assembled a case history for each from upward of 1,500 biographical sources, and had raters estimate a childhood IQ from the documented ages at which each person reached intellectual milestones.
The scale of the effort is not in doubt. Three hundred and one case studies survived, eleven having been discarded for lack of data, and the assembled material ran to roughly 8,500 typed pages. Six raters worked on the estimates, Terman and Maud Merrill and Florence Goodenough among them. Each person received two figures: an AI estimate covering development to the seventeenth birthday and an AII estimate covering seventeen to twenty-six. Each figure carried a coefficient recording how reliable the underlying biographical record was.
That last column is where the method turns on itself. Cox found, and reported plainly, that the more reliable the data, the higher the resulting estimate. She called the relationship entirely unexpected and noted it was discovered only after all the ratings had been tabulated. In other words, a large part of what the numbers track is how much paper survived about a childhood, not how able the child was. Her correlation between the AII estimate and rank order of eminence was .25, and it fell to .16 once the reliability of the data was held constant.
What eminence meant in that correlationCattell's rank order was built from the amount of space each figure occupied in biographical dictionaries. Rank 1 took more column inches than anyone else in the group. So the .16 residual correlation relates an estimate made from surviving records to a measure of how much was written about the person, inside a group already selected for fame.
Both consequences of that drop are worth stating plainly. Part of the apparent association between estimated childhood ability and adult eminence was an artifact of the archive: the more famous the figure, the more survives about the childhood, and the more that survives, the higher the estimate her raters produced. What remains after the adjustment is small. A residual of .16 inside a group already selected for fame means the estimates barely ordered the subjects on the outcome the volume existed to explain, and Cox described it in exactly those terms, as a very small but positive correlation.
Her correction procedure makes that dependency explicit rather than hiding it. Because thin records depressed every estimate, she adjusted each group upward in proportion to the unreliability of its data, and concluded that the corrected mean for the whole group was probably not below 165 against an obtained AII mean of 145. The subgroup figures show the adjustment doing its work: soldiers, whose childhood records were the thinnest in the study, moved from an obtained mean of 125 to a corrected 140, while philosophers moved from 156 to 180. Cox was careful about how far that could be pushed, writing that the purpose of the corrected estimates was not to furnish exact values for an individual or even for a group, but to indicate the region of the scale in which the true figure probably lay.
Cox herself warned against the reading that made the volume famous. Discussing two of her cases, she wrote that a difference of forty-five points between their estimates should not be taken to mean one was that much brighter than the other, because each score simply registers the average rating of everyone who does the things recorded of that subject at that age. The estimated figures still circulate detached from every one of those qualifications, which is why the celebrity number pages on this site, including Einstein and Leonardo da Vinci, trace each claim to its source instead of repeating it. Running that trace across thirty five famous cases leaves four with any cognitive administration on record and exactly one with a named instrument, a date and a comparison sample tested alongside him, which is the tally kept in the audit sorting every celebrity figure by evidence class.
10 What a Very High Score Can and Cannot Support
Precision falls away at the top of any scale, and the reason is structural rather than a flaw in one instrument. A test has fewer items that discriminate at the extremes than in the middle, because items that only a small fraction of people pass contribute little information about everyone else. Fewer discriminating items at the ceiling means a single lucky or unlucky item moves the score further, and the norm sample contains fewer people at that end to anchor the conversion.
That is why the distance between 140 and 145 is not comparable to the distance between 100 and 105 in terms of measurement confidence, even though both are five points. It is also why any responsible report prints a confidence interval alongside the score and why classification bands at the top are wider than bands in the middle. The mechanics are set out on the reliability and validity page, and the score to percentile relationship at the extremes is on the score versus percentile page.
For ACIS specifically, the published figures are these. The Full Scale IQ composite has an omega reliability of .9886 and a g loading of .958, with a standard error of measurement of about 1.60 IQ points. The higher order g confirmatory model fits at CFI .9761, TLI .9726, RMSEA .0406 and SRMR .0217, with a chi-square of 916.703 on 166 degrees of freedom. The technical analysis set for those tables is 2,750 complete records, and the adult reference frame is drawn from 3,243 English speaking records aged 16 to 90. The full derivation is in the technical manual.
The limitation that travels with every one of those numbersACIS is an unsupervised online assessment, not a clinical instrument, and its reference frame is a modelled adult frame built from self-selected respondents rather than a census sample. A reliability coefficient describes consistency within that frame. It does not convert an unsupervised session into a proctored one, and no ACIS result should be used for diagnosis, hiring, accommodations or society admission.
None of this makes a high score meaningless. It means the honest object is a band with an interval around it, not a point. What that band implies in terms of rarity is on the rarity calculator, and what one specific upper tail figure does and does not support is worked through on the score page for 145.
11 The Question Worth Asking Instead of the Threshold
A single number cannot answer the question people are usually asking when they search for a genius threshold. The underlying question is almost always about a specific capability: whether someone reasons well with unfamiliar material, whether they hold and manipulate more in mind than most, whether their verbal knowledge outruns their spatial work. A composite averages all of that into one figure and discards the shape.
The structure that recovers the shape is the CHC framework, which organizes ability into broad domains rather than a single quantity. ACIS reports six: Gf fluid reasoning, Gc comprehension knowledge, Gq quantitative reasoning, Gv visual spatial processing, Gwm working memory and Gs processing speed, expressed as the VCI, FRI, QRI, VSI, WMI and PSI indices. Twenty subtests feed them, with scaled scores on a mean of 10 and a standard deviation of 3, and composites on the familiar mean of 100 and standard deviation of 15. The framework is explained on the CHC model page and the domain by domain breakdown is on the cognitive domains page.
Profiles at the top are rarely flat, and that is expected rather than alarming. A person can sit two standard deviations above the mean on Gc and near the middle on Gs, and the composite that results describes neither accurately on its own. Dispersion across the six indices is the normal case at high ability levels rather than a warning sign. Where a profile is uniformly strong, the composite carries nearly all of the information. Where it is uneven, the composite still describes the overall level accurately while the index pattern supplies direction the single figure cannot: which task formats are handled fastest, where accumulated knowledge outruns reasoning on unfamiliar material, whether working memory is functioning as a support or a constraint. Neither reading undermines the other. A composite of 142 built from a flat profile and a composite of 142 built from a Gc peak against a Gs trough describe the same standing and different working machinery, and only the second description tells anyone what to look at next. Which is why the index pattern carries information the Full Scale figure cannot, a point developed on the fluid versus crystallized comparison.
If the interest is in what the composite is actually built from, the page on what IQ measures and the g factor explanation cover the statistical core, including the evidence behind treating a general factor as the backbone of the composite and the task formats that load on it.
The distinction that matters is not which label sits at the top of the table but whether the administration was supervised. A proctored session with a licensed psychologist using a current battery produces a result that institutions accept, because the conditions were controlled, the identity was verified, and a qualified examiner watched the performance. That is what a school, a court or a clinical process needs, and nothing administered over the internet substitutes for it. The routes and costs are laid out on the where to take a test page and the professional versus online comparison.
An unsupervised online assessment answers a different question, and is worth taking on its own terms rather than as a cheaper version of the first thing. It gives a person a structured estimate of their own profile, at a level of psychometric care that varies enormously across the market. The differences that separate a normed battery from a quiz are set out on the free versus validated comparison.
ACIS is the second kind. It runs 20 subtests across the six CHC domains, reports a Full Scale IQ with six primary indices, and is built for adults aged 16 to 90 in three tiers: Quick at 15 dollars, Optimized at 30 dollars and Full Scale at 50 dollars. It is unsupervised, it is not a clinical instrument, and it is not appropriate for diagnosis, hiring decisions, accommodation requests or high IQ society admission. What the session itself involves is described on the adult test page and the online test page, with the case for a longer battery on the accuracy page.
For research, screening and educational work there is a separate route. ACIS Professional is a workspace with private participant links, no participant personal data, verified scores and CSV export, with a free account that includes three Quick administrations and pay as you go pricing after that. Details are on the research landing page and in the administration section of the home page.
Above a high enough score, the number stops ordering achievement, and the study designed to prove otherwise is what shows it. That is the defensible claim on this page and it rests on data rather than sentiment. Terman selected a cohort at the top of a childhood distribution and followed it for decades. Within that cohort, the men judged most and least successful differed by five points of childhood Stanford-Binet in 1947 and by 7.6 points in the 1968 reselection, while differing by a factor of four in income and by more than fifty percentage points in college completion.
Cox's historical estimates point the same way from a different direction. Inside a group already selected for eminence, her childhood estimates correlated .16 with rank order of eminence once the completeness of the biographical record was held constant. Two independent attempts to make a score predict extraordinary achievement, one prospective and one retrospective, produced small relationships inside the selected range. That is what restriction of range looks like when the selection is on the predictor.
The correct inference is narrow. It is not that the score is meaningless, since the cohort averages on education, occupation and health came out clearly above their contemporaries. It is that once everyone in a comparison is far above the mean, the remaining variance in the score explains little of the remaining variance in what they do. A threshold cannot carry the weight that the word genius puts on it, and the number 140 was never load bearing to begin with. It was a suggestion in a 1916 table that its author retired in 1937.
The professional framework governing all of this is explicit. The International Test Commission guidelines on test use and the Standards for Educational and Psychological Testing (2014), published jointly by the American Educational Research Association, the American Psychological Association and the National Council on Measurement in Education, require that score interpretation be tied to documented evidence for the specific use, that classification labels be reported with their measurement error, and that the limits of the norm sample be disclosed. APA testing standards make the same demand of anyone who reports a classification: state what the number supports, state what it does not, and do not let a label do work the evidence cannot back.
14 Frequently Asked Questions
What IQ is considered genius?
There is no current answer, because no test in professional use has a genius band. The figure people cite, 140, comes from a suggested classification table in Terman's 1916 Stanford-Binet manual, and Terman removed the word from his own 1937 revision.
Did Terman ever define genius by a score?
He explicitly declined to. In the 1916 manual he wrote that intelligence tests had not been in use long enough to allow genius to be defined in terms of IQ, and he described the boundaries in his own table as absolutely arbitrary.
What does the WAIS-IV print at the top of its classification table?
Very Superior. Pearson's published WAIS-IV sample score report shows a Full Scale IQ of 139 described as Very Superior in the Qualitative Description column, and the word genius does not appear in the document.
What does the WAIS-5 use instead?
Directional descriptors. The 2024 sample score report labels an index of 124 as Very high and a Full Scale IQ of 111 as Above average. Per the WAIS-5 manual, the top band at 130 and above is Extremely High.
Why did publishers move away from words like superior?
A qualitative descriptor is read aloud to a person who has just been tested, so the vocabulary has shifted toward describing where a score falls rather than what kind of person produced it. That change is visible across the two Pearson reports.
Is a 1916 IQ of 140 the same as a modern 140?
No. The 1916 scale produced a ratio score, mental age divided by chronological age times 100, on a small California standardization group. Modern scores are deviation scores ranked against an age matched national reference sample.
Who were the Termites?
The 1,528 California schoolchildren, 856 boys and 672 girls, whom Terman selected from 1921 onward for his longitudinal study and followed for the rest of their lives. They were chosen at roughly 135 and above, with 140 as the target threshold.
Did any of them win a Nobel Prize?
None did. Two boys screened during selection and turned away later won the Nobel Prize in Physics: William Shockley in 1956 and Luis Alvarez in 1968.
Does that prove IQ testing does not work?
No, and the researchers who studied it most carefully say so. Warne, Larsen and Clark reported in Intelligence in 2020 that a simulation of Terman's procedure gave a 53 to 83 percent chance of catching neither laureate, purely from base rates.
How did the Terman cohort turn out overall?
As a group they finished more education, entered professional occupations more often, earned more and reported better health than their contemporaries. The threshold predicted the ordinary run of good outcomes well and rare eminence poorly.
What were the A and C groups?
In 1940 three judges reviewed 730 men from the cohort aged twenty-five and over and identified the 150 most successful, labelled A, and the 150 least successful, labelled C, roughly the top and bottom fifth by vocational achievement.
How far apart were the A and C groups on childhood IQ?
Five points. Terman and Oden reported means of 155.0 and 150.0 on the 1922 Stanford-Binet in Table 100 of the 1947 volume, with the Terman Group Test showing no reliable difference at all.
What did Oden find twenty years later?
Reselecting 100 men at each end for the 1968 report, she found childhood Stanford-Binet means of 157.3 and 149.7. She noted that 17 percent of the least successful group scored above the most successful group's mean.
How large was the outcome gap between those two groups?
Very large. Median total family income in 1959 was 33,125 dollars for the A group against 8,500 dollars for the C group, and 92 percent of the A men had graduated from college against 40 percent of the C men.
What did Terman conclude from that comparison?
That the difference had to lie elsewhere. His summary of the table reads that where all are so intelligent, differences in success must be due largely to nonintellectual factors, and he wrote that intellect and achievement are far from perfectly correlated.
What was Cox 1926 and how did it work?
Catharine Cox built case histories for 301 eminent historical figures from more than 1,500 biographical sources, roughly 8,500 typed pages, and had six raters estimate a childhood IQ for each from documented developmental milestones.
What is the main problem with Cox's estimates?
She discovered, after tabulating everything, that the better the surviving biographical record, the higher the estimate. Part of what the numbers measure is how much documentation survived a childhood rather than the ability behind it.
How strongly did Cox's estimates predict eminence?
Weakly. The correlation with Cattell's rank order of eminence was .25, and it dropped to .16 once the reliability of the underlying data was held constant, inside a group already selected for fame.
Why are scores less precise at the top of the scale?
Tests carry fewer items that discriminate among very high performers, and norm samples contain fewer people at the extremes to anchor the conversion. Both effects widen the interval around a high score relative to a mid range one.
Does ACIS use a genius band?
No. Its public table runs Average, High Average and Superior, then four gifted bands from Moderately Gifted at 130 up to Profoundly Gifted at the top. The word genius appears nowhere in the classification.
Has ACIS assessed any of the people discussed here?
No. Terman's cohort members and Cox's historical figures were never assessed with this instrument, and no retrospective assessment is possible. Everything above evaluates published records rather than producing new estimates.
Take the assessment
You get a profile, not a number
ACIS measures six CHC domains across 20 subtests and reports each one with its own normed score and confidence interval, so you can see where you are strong and where you are not.