The hardest IQ test and the arithmetic of difficulty
Difficulty is a property of an item, not a virtue of a test. A question that almost everyone gets wrong tells you almost nothing about the people who got it wrong, and a battery assembled entirely from such questions measures a narrow slice of the population badly. This page is about where difficulty comes from, how ceilings get built, and why the hardest test and the most precise test are almost never the same instrument.
Frustration is not evidence of measurement. An item that everybody fails contributes nothing to a score, which is the central arithmetic of this page.
0 Quick Answer
There is no single hardest IQ test, and the tests that advertise themselves that way are usually measuring less precisely than the ones that do not. Difficulty in psychometrics is a number attached to an item, not a badge attached to a battery. In classical test theory it is the proportion of a sample that answers the item correctly, which Hambleton and Jones set out in their 1993 NCME instructional module in Educational Measurement: Issues and Practice, 12(3), 38 to 47. In item response theory it is the b parameter, the point on the ability scale where the probability of a correct answer crosses one half.
Once difficulty is a number rather than an adjective, the interesting question changes. An item measures best near its own difficulty and contributes very little far away from it. Hambleton and Jones state it plainly: items "tend to make their best contribution to measurement precision around their b value on the ability scale." A test made entirely of very hard items therefore concentrates all of its measurement power at the top and leaves the middle of the distribution close to unmeasured. That is a design choice with a cost, not a free upgrade.
The phrase itself has a documented origin. The April 1979 cover of Omni magazine carried the line "The world's hardest I.Q. test, complete text," and the contents page credits the article on page 116 to Kevin Langdon. You can read the scanned issue hosted by the Mega Society. What followed over the next four decades is a useful natural experiment in what happens when a test chases difficulty without the machinery that makes a score comparable.
57 percent
Of the variance in mean error rates across 32 Raven problems explained by one variable, the number of rules in the problem, in Carpenter, Just and Shell's 1990 analysis.
16 of 2,200
Children in the WISC-V normative sample who obtained any composite score of 150 or higher, reported by Pearson in its 2019 extended norms technical report.
Top 2 percent
American Mensa's published admission criterion, which no unsupervised or internet administered test can satisfy under its own stated rules.
Item difficulty is measured, not designed, and the best documented account of what drives it in a reasoning test comes from an eye tracking study published in 1990. Patricia Carpenter, Marcel Just, and Peter Shell of Carnegie Mellon University published What One Intelligence Test Measures in Psychological Review, 97(3), 404 to 431. They took the Raven Progressive Matrices apart item by item, using verbal protocols, eye fixation patterns, and error analysis from college student samples of 12 and 22 people, and then built two computer simulations that reproduced the performance of the median and the best performers.
Their first finding was that almost every problem in the test can be described by five rule types. Constant in a row, where the same value occurs across a row but changes down a column. Quantitative pairwise progression, where an attribute increases or decreases between adjacent entries. Figure addition, where the element in the first column is superimposed on the element in the second to produce the third. Distribution of three values, where three values of an attribute are spread across a row. Distribution of two values, where two values are spread across a row and the third position is null.
Their second finding is the one that matters here. Difficulty was not driven by exotic content. It was driven by how many rules a problem required at once. In their words, a simple linear regression whose single independent variable was the total number of rules in a problem, not counting constant rules, "accounted for 57% of the variance among the mean error rates" across the 32 problems they had classified. Their error rates correlated .91 with the error rates Forbes had reported in 1964 from 2,256 British adults, so this was not an artifact of a small student sample.
The mechanism they proposed is more useful than the number. Extra rules do not make the individual inferences harder. They make the problem harder to keep track of. Carpenter, Just, and Shell concluded that the processes distinguishing individuals were "primarily the ability to induce abstract relations and the ability to manage dynamically a large set of problem-solving goals in working memory." Hardness, in a well built reasoning item, is load. It is not obscurity, trick phrasing, or a wider vocabulary.
Why this rules out a whole category of hard questionAn item can be made unanswerable in a dozen cheap ways: an ambiguous stem, an answer key that depends on a convention the test taker was never told, a required fact that only a specialist would hold. Every one of those raises the error rate without measuring reasoning. The Carpenter, Just, and Shell result is valuable precisely because it identifies a source of difficulty that scales with the construct rather than with the test taker's luck.
2 A Hard Item and a Discriminating Item Are Different Things
An item that everybody fails discriminates nothing, and this is arithmetic rather than opinion. In classical test theory the difficulty index is the proportion of the sample answering correctly, written p, and item discrimination is usually a point biserial correlation between the item and the total score. Hambleton and Jones name both in the 1993 module: "classical item statistics such as item difficulty (i.e., proportion correct) and item discrimination (i.e., point biserial correlations)."
Now push p to its limits. A binary item's variance is p times one minus p. At p equal to one, everybody passes and the variance is zero. At p equal to zero, everybody fails and the variance is zero again. A variable with no variance cannot correlate with anything, so the item's discrimination is undefined, and the item contributes exactly nothing to the ranking of the people who took it. It occupies testing time and returns no information. The maximum of that variance sits at p equal to one half, which is why item writers who want information rather than intimidation aim at items that split the target group.
This is the single most common failure of tests marketed on difficulty. If a test is built so that the modal score is near zero, most of its items are at p near zero for most of its takers, and the score is being carried by a handful of items that happened to be answerable. The test feels hard. It measures a fraction of what its length suggests.
Item property
What it means
What it does for the score
Where it belongs
p near .90
Almost everyone passes
Almost no ranking information, but establishes the floor and lets low scorers register a result
Early positions in a subtest
p near .50
The group splits evenly
Maximum item variance and the best chance of a high point biserial
The body of the subtest
p near .05
Almost everyone fails
Ranks only the small group who can pass it, which is useful only if that group is the target
The ceiling, in small numbers
p equal to 0
Nobody in the sample passes
Zero variance, zero discrimination, zero contribution
Nowhere, once the data comes back
High p with negative discrimination
Easy but the wrong people pass it
Actively degrades the total score
Removed, usually a keying or wording fault
The last row is worth a note. A negative point biserial means the stronger test takers are getting the item wrong more often than the weaker ones, which almost always signals a broken key or an ambiguous stem rather than a genuine reversal of ability. Difficulty statistics are also sample dependent, a limitation Hambleton and Jones state directly: classical item statistics "are dependent on the examinee sample in which they are obtained." A test normed on a self selected group of puzzle enthusiasts has p values that describe puzzle enthusiasts and nobody else, which is why the same statistics behave differently once a real norm sample is involved.
3 Where a Test Puts Its Information
Item response theory replaces the single difficulty index with a location on the ability scale, and once you have that, you can say exactly where a test is precise and where it is not. The b parameter is the point on the ability continuum at which the probability of a correct answer reaches one half in the one and two parameter models. Nguyen, Han, Kim, and Chan define it in The Patient, 7(1), 23 to 35, 2014, as the "point on the trait scale where the probability of endorsing an item is 50%," and the a parameter as the slope of the item characteristic curve at b.
Each item then carries an information function. Nguyen and colleagues state that "item information is typically highest in the region of the trait near the location parameter, b; items with greater discrimination contribute more information." Sum those curves across a test and you get the test information function, and Hambleton and Jones give the relationship that turns it into something a reader can use. The standard error of the ability estimate is one divided by the square root of the information at that point, so "the more information provided by a test at a particular ability level, the smaller the errors associated with ability estimation."
That single equation explains almost everything about hard tests. Information is a finite budget. Spend it at the top and the middle goes cheap. Spend it in the middle and the top goes cheap. There is no arrangement of items that makes a test maximally precise everywhere at once, because each item can only sit at one place on the scale.
SE equals one over the square root of information
The relationship given by Hambleton and Jones in the 1993 NCME module, which is why measurement error is a function of ability level rather than a single number for the whole test.
Maximum near b
Where an item contributes most, per Nguyen et al. 2014 in The Patient. An item three standard deviations above a test taker tells you almost nothing about them.
Discrimination scales it
Two items at the same difficulty can differ several fold in information, because the a parameter multiplies the height of the curve.
The practical consequence for a general purpose battery is that its items should be spread across the range it claims to measure, weighted toward where its users actually sit. A test that reports scores from 40 to 160 and places every item above the 99th percentile is not a harder version of a normal test. It is a different instrument with a different target, and its score in the average range is close to meaningless. The relationship between a score, its error band, and the population is developed further on the reliability and validity page and on the page on the 15 point standard deviation.
4 How a Ceiling Is Actually Built
A test ceiling is not a policy decision, it is a table, and the clearest public document showing how one gets extended is Pearson's own. In November 2019 Susan Engi Raiford, Troy Courville, Daniel Peters, Barbara Gilman, and Linda Silverman published WISC-V Technical Report Number 6, Extended Norms. It exists, in the authors' words, because of "requests from the National Association for Gifted Children," and it is designed "to more clearly identify highly gifted children with composite scores far above 130."
The mechanism is instructive. An ordinary Wechsler subtest converts a range of raw scores into a scaled score of 19, the published maximum. The extended tables reach into that range and split it further, so that raw scores which previously all mapped to 19 now map to 20, 21, and upward to a maximum of 28, with composites reaching 210. Nothing new is added to the test. The headroom was always in the item set, and the published table simply stopped using it.
Which is exactly why the extension has limits, and the report says so. "Not all subtests have extensive room for extension in this range, especially at older ages, and extending the scaled score range to 28 is not always possible." The tables in that report contain dashes at scaled score positions that no attainable raw score reaches. A scale position that no performance can produce is not a measurement, and the authors decline to invent one.
The rarity figures in the same report show why this territory is thin. Among the 2,200 cases in the WISC-V normative sample, only 16 children obtained any composite score of 150 or higher. The authors add the projection directly: "out of 20,000,000 same-age peers, only one child would be expected to obtain an FSIQ of 180 or higher." The validity sample they eventually assembled from gifted centers was 108 usable cases, mean age 9.6 years, and 43 percent of those children saw an increased Full Scale IQ under the extended tables.
The detail that undercuts most ceiling claimsIn that same highly gifted sample, the extended Verbal Comprehension Index reached a maximum of 180 while the extended Processing Speed Index reached only 148. The extension is bounded by how many items a subtest has left above the ordinary ceiling, and speeded subtests run out first. A battery's advertised top number is a property of its longest tailed subtests, not of the battery as a whole.
5 Tests by Ceiling and What Each Ceiling Rests On
Ceilings are comparable only when you also state what is holding them up. Two instruments can both print a number near 160 and mean entirely different things by it, because one is reporting a rank against a documented sample and the other is reporting an extrapolation from a curve. The table below sorts the instruments people ask about by what their top figure rests on. Why standard clinical batteries stop where they do is the subject of the high range page rather than this one.
Instrument
Top of the reported range
What the ceiling rests on
What it will not support
Wechsler adult scales, standard tables
Around 160 composite
Rank against a stratified norm sample that contains very few cases at the extreme
Distinctions above the published cap, which the tables simply do not print
WISC-V with extended norms, Pearson 2019
Composite 210, subtest scaled score 28
Unused raw score headroom in the existing item set, plus a 108 case gifted validity sample
Equal extension across subtests, since the report states extension to 28 is not always possible
Raven's Advanced Progressive Matrices
Raw score out of 36 in Set II, reported as a percentile against the reference group used
An item set written for adults and adolescents of above average ability, 48 items across a 12 item Set I and a 36 item Set II
A composite IQ, a domain profile, or any verbal or working memory reading
The Mega Test, Hoeflin, published 1985
Claimed to reach a rarity of one in a million
Equipercentile equating against scores that respondents reported about themselves
A defensible percentile, because the reference group selected itself and the anchor scores were unverified
The Langdon Adult Intelligence Test, 1979
Accepted by the Mega Society at IQ 175 for results before 1994
A 553 person norming fitted to self reported prior test scores
Comparability with a supervised administration, which its own author never claimed
ACIS Full Scale
Composite reported with its measurement error, six primary indices from 20 subtests
Adult reference frame
Clinical diagnosis, hiring decisions, accommodations, or high IQ society admission
Read the third column rather than the second. A number is only as portable as the reference group behind it, and three of the six rows above rest on reference groups that were assembled from people who volunteered because they expected to do well. That is not a small technical caveat. It is the difference between a percentile and a compliment. What Raven's 2 covers in its current edition is set out on the Raven's 2 page, and the wider inventory of instruments in current use is on the types of tests page.
6 The Omni Era and the First World's Hardest Test
The modern idea of a hardest IQ test entered general circulation through a science magazine, and the primary documents are still readable. The April 1979 issue of Omni, volume 1 number 7, ran the cover line "The world's hardest I.Q. test, complete text." The contents page lists it as an article by Kevin Langdon beginning on page 116. Langdon, working outside the test publishing industry, had written a set of items intended to reach far above the range of ordinary instruments, and the magazine printed it in full with instructions for readers to mail their answers in for scoring.
What happened next is documented in Langdon's own norming report. His second norming study of the Langdon Adult Intelligence Test, dated 15 July 1979 and archived by Darryl Miyaguchi with Langdon's permission, covers 553 testees, 207 on Form A and 346 on Form B. The report states that "only a handful of the earliest responses to the test's appearance in the April 1979 issue of Omni are included," so the numbers describe the pre Omni respondents rather than the magazine's readership.
The internal consistency figures were respectable. Langdon reports Spearman Brown corrected split half correlations of .902 for Form A and .898 for Form B. Reliability was never the weakness. The weakness was the anchor. Langdon calibrated LAIT scores against scores that respondents reported having obtained on other tests, a list that in his own table includes the Stanford-Binet, the Army General Classification Test, the Miller Analogies, the Wechsler Adult Intelligence Scale, the SAT, the GRE, and Cattell's scales.
Langdon said so himself, twenty one years later. In an introduction dated 22 August 2000 he wrote that the second norming "was done with what I now regard as very crude statistical methods," and described "many outlying points, both super-high scores and scores from outside the main, self-selected population taking the test." That is an unusually honest disclosure from a test author, and it is the correct place to start any evaluation of the high range test tradition: the internal reliability was real, the external anchor was self reported, and the population was self selected.
7 The Mega Test, the Titan Test, and the Norming Problem
Six years later the same magazine ran the test that became the genre's landmark, and the peer reviewed re-analysis of it did not arrive until 2020. Scot Morris's article "The one-in-a-million I.Q. test" appeared in Omni in April 1985 at pages 128 to 132, presenting Ronald K. Hoeflin's Mega Test: 48 questions, half verbal and half mathematical, untimed and unsupervised, taken at home over as long as the reader wanted. Hoeflin followed it with the Titan Test, published in Omni in April 1990 under the headline "Mind Games: the hardest IQ test you'll ever love suffering through," at page 90 and following.
Both were normed the way the LAIT had been, by lining raw scores up against scores respondents said they had obtained elsewhere. Nothing in that procedure verifies an anchor score, and nothing in it corrects for the fact that the people who mail in answers to a magazine puzzle test are not a sample of anything. The tests nonetheless became the admission route for the most selective high IQ societies, and their scores circulated for three decades as though they were IQ points.
The first mainstream academic scrutiny was published in 2020. David Redvaldsen of the University of Agder published Do the Mega and Titan Tests Yield Accurate Results? in Psych, 2(2), 97 to 113, DOI 10.3390/psych2020010. He renormed both tests against the normal curve and reached a mixed verdict that is more interesting than a dismissal. His finding was that "although official scores reported to test-takers are too high, it is likely that the Mega Test does stretch to the one in a million level. The Titan Test does not."
Three numbers from that paper are worth carrying. Test takers who had previously sat standard intelligence tests reported average scores of 135 to 145 on them. The mean of all scores on the Mega Test was IQ 137 and on the Titan Test IQ 138. Redvaldsen reads this as evidence that respondents "had considerable scope to find their true level without ceiling effects," which is the correct technical reading and also a complete description of the sample: a group already two to three standard deviations above the mean, measured against itself.
What the hardest test in the genre was criticized forRedvaldsen's judgement of the Mega Test is that it "has a higher ceiling and a lower floor than the Titan Test" and that it "is somewhat marred by a number of easy items in its verbal section." The instrument famous for difficulty was faulted, in the only peer reviewed analysis of it, partly for containing items that were too easy. Difficulty is a distribution across an item set, not a slogan.
Both tests have since been retired by the society that used them. The Mega Society's own admissions page lists the Mega Test as qualifying only for scores obtained before 1995 and the Titan Test only before 1 September 2020, in each case because the test was compromised. Answers that circulate cannot be unpublished, and an untimed take home test has no defense against that.
8 Why an Untimed, Unproctored, Self Selected Test Cannot Produce a Defensible Percentile
A percentile is a claim about a population, and three separate features of the high range test format each break that claim on their own. They are worth separating, because they are usually discussed as one vague objection when they are three specific and independent ones.
The reference group selected itself. A percentile says what fraction of a defined population scores below you. The people who mailed answers to Omni chose to, because they expected to do well. Scoring against that group understates a given performance, which then invites the author to shift the scale upward to compensate, and the resulting numbers have no fixed relationship to the population everybody else is measured against. This is the same failure mode described on the norming page, applied at the extreme.
The anchor scores were unverified self reports. Equating a new test to old ones requires the old scores to be real. In the LAIT and Mega norming procedures, respondents supplied their own prior figures. Errors of memory, of edition, and of self presentation all push in the same direction, and none of them are detectable from the answer sheet.
The administration was uncontrolled. Untimed and unsupervised means unbounded time, reference material, collaboration, and repeat attempts are all available and none are recorded. A raw score from such a session is a joint measure of ability and of how much effort and how many resources the taker chose to apply, and the two cannot be separated after the fact.
The societies that use these instruments are explicit about the tradeoff rather than evasive. The Mega Society states on its own site that it uses "untimed, unsupervised IQ tests that have been normalized using standard statistical methods" precisely because it holds that supervised timed instruments cannot discriminate at the one in a million level it admits at. That is a coherent position. It is also an admission that the price of reaching that far is giving up the conditions that make a score comparable.
Every one of these problems is independent of item quality. A high range test can contain excellent items, as Carpenter, Just, and Shell showed the Raven's does, and still produce an uninterpretable number, because the number is a function of the norming and the administration rather than of the items. What separates a validated battery from a puzzle set is set out on the free versus validated comparison, and the conditions that separate supervised from unsupervised results are on the professional versus online page.
9 What a Hard Score Actually Buys You
No institution accepts a score because the test was hard, and the published criteria show what they do accept instead. American Mensa's requirement is a result in the top two percent of the general population. Its qualifying test scores page lists the specific figures: a Full Scale IQ of 130 on the Wechsler scales, 130 on the Stanford-Binet 5, 132 on the older Stanford-Binet, 132 on the Otis Lennon composite, 148 on the Cattell, and 131 on the fourth edition of the Woodcock-Johnson.
The same page carries the sentence that settles the question for anything taken at home. American Mensa "does not accept unsupervised testing as proof of eligibility, specifically unsupervised testing administered electronically or via Internet-based tests," and requires that tests "be administered by a neutral and qualified third party in a traditional testing environment." No amount of difficulty substitutes for that. A supervised 130 qualifies and an unsupervised 175 does not, which tells you exactly what the institution is buying: controlled conditions, verified identity, and a documented reference sample.
What you want
What actually satisfies it
What difficulty contributes
Society admission
A supervised administration of an instrument on the society's approved list, at or above its published cutoff
Nothing. The cutoff is a percentile, and the approved list is short
A clinical or educational record
A licensed examiner, a current battery, a written report, and documented conditions
Nothing. Difficulty is not one of the acceptance criteria
A defensible personal estimate
A normed battery with enough subtests to produce a stable composite and a stated error band
Some, but only in the form of items placed where the test taker actually sits
A profile that explains an uneven result
Multiple domains measured separately, each with its own reliability
None directly. Breadth does this work, not difficulty
A number that is larger than the last one
An untimed self scored test with self selected norms
All of it, and nothing else
The last row is included because it is the honest description of a real demand. Wanting a bigger number is not shameful and it is not rare. It is simply a different want from wanting a measurement, and the instruments that serve it are not the instruments that serve the other four rows. The published admission criteria across the societies are collected on the society requirements page, and what Mensa's own supervised testing session involves is on the Mensa page.
10 Which Tasks Carry the Most Information in a Modern Battery
If the goal is measurement rather than difficulty, the question to ask about a task is how much of the general factor it carries, not how many people it defeats. A g loading is the correlation between a subtest and the general factor extracted across the whole battery. It is a statement about how much a task tells you about overall ability, and it is largely independent of how hard the task feels.
In the ACIS technical manual's normative model, the highest loading reasoning subtests are Logic Grid at .884, Figure Weights at .872, and Complex Relations at .856, with Matrix Reasoning at .840 and Visual Number Series at .837. The quantitative subtests sit alongside them, with Mathematical Achievement at .908 and Arithmetic at .897. At the other end, Symbol Search loads .752 and Visual Sequence .738. The full derivation is in the technical manual.
Notice what that ordering does not track. Logic Grid presents a constraint satisfaction problem where several conditions have to be held together and combined, and Figure Weights asks for a quantitative inference across balance displays. Both carry more of the general factor than Complex Relations, which most test takers find harder to describe afterward. Subjective difficulty and psychometric information are not the same axis, and confusing the two is how test design goes wrong.
.922
The g loading of the ACIS Fluid Reasoning Index, the highest of the six primary indices in the published technical manual, built from five subtests.
.954
The g loading of the General Ability Index, drawn from 15 non speeded reasoning and knowledge subtests, with a standard error of measurement of 1.61 points.
.648
The g loading of the Processing Speed Index, the lowest of the six, which is why a two subtest speed composite carries a standard error of 5.22 points.
The index level figures make the same point at a larger scale. The Fluid Reasoning Index reaches an omega of .9727 across five subtests with a standard error of measurement of 2.48 points, while the Processing Speed Index reaches .8790 across two subtests with a standard error of 5.22. More indicators buy precision. Harder indicators do not, unless they happen to sit where the test taker sits. What the fluid reasoning tasks are and what they measure is set out on the fluid reasoning test page, and the statistical core of the general factor is on the g factor page.
11 What Happens to Measurement at the Top of a Scale
Precision is not constant across a score range, and the reason is the same information budget from earlier applied to a region where both items and people are scarce. Two things thin out together at the extreme. There are fewer items whose b parameter sits that high, so the test information function falls, and by the relationship Hambleton and Jones give, the standard error rises as the square root of that falling information. Separately there are fewer people in the reference sample at that level to anchor the conversion, which is the point Pearson's own figure of 16 cases in 150 or above out of 2,200 makes concrete.
The consequence for a reader is that a five point difference does not mean the same thing everywhere. Near the middle of the distribution the norm table is dense and the items are plentiful. Near the top the same five points may be one raw item, converted through a table row that rests on a handful of cases. A responsible report handles this by printing a confidence interval rather than a point, and by widening the classification bands at the extremes.
The ACIS figures follow that pattern. The General Ability Index, built from 15 subtests, carries a standard error of measurement of 1.61 points and an omega of .9885. The Quantitative Reasoning Index, built from two, carries 4.02 points and .9283. The Processing Speed Index, also from two, carries 5.22 and .8790. Those are population level standard errors, and they widen further as a score moves away from the region where the model has the most data. How a score maps onto a rank is explained on the score versus percentile page and the underlying scale is on the IQ score scale.
The limitation that travels with every figure on this pageACIS is a self-administered online assessment, not a clinical instrument. Its adult reference frame is documented in the technical manual. A reliability coefficient describes consistency within that frame and does not convert a self-administered session into a proctored one. No ACIS result should be used for diagnosis, hiring, accommodations, or society admission. Separately, this page could not verify a norming study for the Haselbauer-Dickheiser Test for Exceptional Intelligence in any indexed journal, so it is named here without figures attached.
12 What to Take Instead, and Why
The useful version of the question is not which test is hardest but which test places its items where you are. That is answerable in advance, because it is a design property rather than an experience. A battery that reports a composite plus separate domain indices, from enough subtests that each index has more than two indicators, will give a more precise reading of a strong performer than a short difficult puzzle set, even though the puzzle set will feel more like an achievement.
ACIS is built on that principle. It runs 20 subtests across the six CHC domains for adults aged 16 to 90, in three tiers: Quick at 15 dollars with six subtests, Optimized at 30 dollars with 13, and Full Scale at 50 dollars with all 20. Only the Full Scale form produces the Full Scale IQ with all six primary indices and every composite, which is the form that matters if the question is about the top of a profile rather than a rough placement. Subtests discontinue after three consecutive errors or timeouts, so nobody spends 20 minutes failing items far above their level, which is the same information logic described earlier applied to administration time.
Two things it deliberately does not do. It does not print a number above the range its normative model supports, and it does not offer an untimed take home format for its timed subtests. Both refusals cost headline appeal and buy interpretability. What the individual task formats look like is on the subtest formats page, and the case for a longer battery over a short one is on the accuracy page.
If the underlying interest is in how a raw performance becomes a scaled score in the first place, that machinery is worked through on the page on how IQ is calculated. If the interest is specifically in the rarity of a high result rather than the score itself, the rarity calculator converts a figure into a one in X statement.
Every number above traces to one of the following, and each is linked so a reader can check it rather than take it. Where a claim rests on a document that is not freely readable, the page says so instead of implying otherwise.
Hambleton, R. K., and Jones, R. W. (1993). An NCME Instructional Module on Comparison of Classical Test Theory and Item Response Theory and Their Applications to Test Development. Educational Measurement: Issues and Practice, 12(3), 38 to 47. Source of the proportion correct definition of item difficulty, the statement that classical item statistics are sample dependent, the definition of the b parameter, the claim that items measure best near their own b value, and the relationship between test information and standard error. Full module PDF, hosted by NCME.
Nguyen, T. H., Han, H. R., Kim, M. T., and Chan, K. S. (2014). An Introduction to Item Response Theory for Patient-Reported Outcome Measurement. The Patient, 7(1), 23 to 35. Source of the fifty percent probability definition of b, the slope definition of a, and the statement that item information peaks near b. Open access at PubMed Central.
Carpenter, P. A., Just, M. A., and Shell, P. (1990). What One Intelligence Test Measures: A Theoretical Account of the Processing in the Raven Progressive Matrices Test. Psychological Review, 97(3), 404 to 431. Source of the five rule taxonomy, the 57 percent of error rate variance explained by rule count, the .91 correlation with the Forbes 1964 data on 2,256 British adults, and the goal management conclusion. Carnegie Mellon technical version at ERIC.
Raiford, S. E., Courville, T., Peters, D., Gilman, B. J., and Silverman, L. (2019). WISC-V Technical Report Number 6: Extended Norms. NCS Pearson. Source of the extension to scaled score 28 and composite 210, the 16 of 2,200 figure, the one in 20,000,000 projection for a Full Scale IQ of 180, the 108 case gifted sample, and the statement that extension is not always possible. Report PDF, published by Pearson.
Redvaldsen, D. (2020). Do the Mega and Titan Tests Yield Accurate Results? An Investigation into Two Experimental Intelligence Tests. Psych, 2(2), 97 to 113, DOI 10.3390/psych2020010. Source of the renorming verdict, the mean scores of 137 and 138, the 135 to 145 prior score range, and the criticism of easy verbal items in the Mega Test. Open access record at the University of Agder.
Langdon, K. (1979). Statistical Report, LAIT Norming Number 2, 15 July 1979, with an author's introduction dated 22 August 2000. Source of the 553 testee sample, the Form A and Form B split, the corrected split half figures of .902 and .898, the list of tests used for equating, and Langdon's own description of the method as very crude. Archived copy of the report.
Omni magazine, April 1979, volume 1 number 7. Cover line "The world's hardest I.Q. test, complete text," with the article credited on the contents page to Kevin Langdon at page 116. Scanned issue pages. The Mega Test appeared in Scot Morris, The one-in-a-million I.Q. test, Omni, April 1985, pages 128 to 132, and the Titan Test in Omni, April 1990, page 90 following. Those two issues were cited from bibliographic records rather than read directly for this page.
American Mensa. Qualifying Test Scores. Source of the top two percent criterion, the instrument by instrument cutoffs, and the refusal of unsupervised and internet administered testing. Published criteria page.
The Mega Society. Tests accepted for admission. Source of the LAIT threshold of IQ 175 before 1994, the Mega Test raw score of 43 before 1995, the Titan Test raw score of 43 before 1 September 2020, the retirement of both as compromised, and the society's own statement about untimed unsupervised instruments. Admissions page.
ACIS technical manual. Source of every ACIS figure quoted above, including the subtest g loadings, the index omegas and standard errors of measurement, and the composition of the six primary indices. Available at the technical manual page.
Two professional frameworks govern how any of this should be reported. The Standards for Educational and Psychological Testing, published jointly in 2014 by the American Educational Research Association, the American Psychological Association, and the National Council on Measurement in Education, require that a score interpretation be supported by evidence for the specific use claimed, that measurement error be reported alongside the score, and that the composition and limits of the norm sample be documented. The APA standards on test use place the same obligation on whoever reports a result: state what the number supports, state what it does not, and do not let the way a test felt stand in for evidence about what it measured.
14 Frequently Asked Questions
What is the hardest IQ test in the world?
There is no agreed answer, because difficulty is a property of items rather than a ranking of tests. The phrase entered circulation with the April 1979 cover of Omni magazine, which advertised Kevin Langdon's test as the world's hardest IQ test in complete text.
What makes an IQ test question hard?
In reasoning items it is mostly load rather than obscurity. Carpenter, Just, and Shell reported in Psychological Review in 1990 that the number of rules in a Raven problem alone explained 57 percent of the variance in mean error rates across the 32 problems they classified.
What is a p value in item analysis?
It is the proportion of a sample that answers an item correctly, so a higher p means an easier item. Hambleton and Jones note in their 1993 NCME module that this statistic depends on the examinee sample it was computed in, which is why it does not transfer between populations.
What is the b parameter in item response theory?
It is the point on the ability scale where the probability of a correct answer reaches one half in the one and two parameter models. Nguyen and colleagues define it that way in The Patient in 2014, and it places the item on the same scale as the person.
Can an item be too hard to be useful?
Yes, and the point where it becomes useless is exact. A binary item's variance is p times one minus p, so an item nobody passes has zero variance, cannot correlate with the total score, and contributes nothing to anyone's ranking.
Why does a test of only hard items measure the middle badly?
Because an item contributes most information near its own difficulty. Hambleton and Jones state that items make their best contribution to precision around their b value, so a set of items all placed at the top leaves the middle of the distribution with almost no information.
Is a harder test more accurate?
Not in general. Accuracy at a given ability level is governed by the test information function at that level, and the standard error is one over the square root of that information. Difficulty helps only if it moves items closer to where the test taker actually sits.
How is a test ceiling actually raised?
By splitting raw score ranges that the ordinary table collapsed into its top scaled score. Pearson's 2019 WISC-V extended norms report raises subtest scaled scores to a maximum of 28 and composites to 210 using the existing item set, with no new items added.
Can any subtest have its ceiling extended?
No. The Pearson extended norms report states that extending the scaled score range to 28 is not always possible, especially at older ages, because a subtest can run out of raw score headroom. In its gifted validity sample the extended Processing Speed Index reached only 148.
How rare is a very high childhood composite?
Rarer than most people assume. Pearson reports that among the 2,200 cases in the WISC-V normative sample only 16 children obtained any composite score of 150 or higher, and that one child in 20,000,000 same age peers would be expected to reach a Full Scale IQ of 180.
What was the Mega Test?
A 48 item untimed test by Ronald K. Hoeflin, half verbal and half mathematical, published by Scot Morris in Omni in April 1985 at pages 128 to 132. Readers completed it at home over as long as they liked and mailed answers in for scoring.
Were the Mega and Titan tests ever independently evaluated?
Once, in a peer reviewed journal. David Redvaldsen renormed both in Psych in 2020 and concluded that although official scores reported to test takers were too high, the Mega Test likely does stretch to the one in a million level while the Titan Test does not.
What was wrong with how those tests were normed?
They were equated against scores that respondents reported about themselves, from a group that had selected itself by choosing to take a magazine puzzle test. Neither the anchor scores nor the reference group can support a percentile claim about the general population.
Did the authors of these tests admit the limitation?
One did, explicitly. In an introduction dated 22 August 2000 Kevin Langdon wrote that his second LAIT norming used what he then regarded as very crude statistical methods, with many outlying points and scores from outside the self selected population taking the test.
Are the Mega and Titan tests still accepted anywhere?
Not for new results. The Mega Society's own admissions page accepts Mega Test scores only from before 1995 and Titan Test scores only from before 1 September 2020, in both cases because the test was compromised once answers circulated.
Does Mensa accept a score from a hard online test?
No. American Mensa states that it does not accept unsupervised testing as proof of eligibility, specifically unsupervised testing administered electronically or via internet based tests, and requires administration by a neutral qualified third party.
What score does Mensa actually require?
A result in the top two percent of the general population. Its published table lists a Wechsler Full Scale IQ of 130, a Stanford-Binet 5 score of 130, 132 on the older Stanford-Binet, and 148 on the Cattell, among others.
How many items does Raven's Advanced Progressive Matrices have?
Forty eight in total, arranged as a 12 item Set I and a 36 item Set II, written for adults and adolescents of above average ability. It reports performance against a reference group rather than producing a domain profile.
Which ACIS subtests carry the most information?
Among the reasoning tasks, Logic Grid loads .884 on the general factor, Figure Weights .872, and Complex Relations .856 in the published technical manual. Those loadings describe how much each task says about overall ability, not how hard it feels.
Why does a score become less precise at the top of the scale?
Two scarcities compound. Fewer items sit that high, so test information falls and the standard error rises with it, and fewer people in the reference sample sit there to anchor the conversion. Both effects widen the interval around a high score.
Is ACIS a hard test?
It is a broad one. It runs 20 subtests across six CHC domains with items spread over the range it reports, and subtests discontinue after three consecutive errors or timeouts. It is self-administered, is not a clinical instrument, and is not accepted for society admission.
Take the assessment
You get a profile, not a number
ACIS measures six CHC domains across 20 subtests and reports each one with its own normed score and confidence interval, so you can see where you are strong and where you are not.