The Dunning-Kruger effect and IQ: the real numbers
Nobody places their own IQ accurately. Across 41 published studies, self-estimates correlate about .33 with measured ability, which leaves the guess wrong by tens of points in both directions. The popular Dunning-Kruger story adds a second error on top of that one, and this page separates what survives from what does not.
The same data, two conventions. Quartile averaging is what makes the classic crossing pattern appear.
0 The Short Answer
You cannot estimate your own IQ with useful precision, and the reason has nothing to do with being unskilled: across 41 published studies and 154 effect sizes, the meta-analysis by Freund and Kasten in Psychological Bulletin in 2012, How smart do you think you are?, put the correlation between self-estimated and psychometrically measured cognitive ability at r = .33. A larger and more recent synthesis lands in the same place. Patzl, Oberleiter and Pietschnig, publishing in Journal of Intelligence in 2024, pooled 278 effect sizes from 115 independent samples across 93 studies and 36,833 participants and reported r = .30, with a 95 percent confidence interval from .27 to .33.
That single figure answers the practical question, and it answers it for everyone. A correlation of .33 accounts for roughly 11 percent of the variance in measured ability. It does not describe a defect concentrated in the bottom quartile. It describes a weak signal running through the whole range, which means the confident person and the modest person are both guessing, and both are usually wrong by a lot.
The popular version of the Dunning-Kruger effect gets this wrong twice. It treats bad self-estimation as a property of low performers, when the data show it is general. And it rests on a chart that a substantial body of methodological work says is partly manufactured by regression to the mean plus the tendency of almost everyone to rate themselves above the midpoint. This page works through both problems with the sources in hand, and it stops short of the fashionable conclusion that the phenomenon is fictional, because a real metacognitive component survives the critique and can be pointed to.
r = .33
Correlation between self-estimated and measured cognitive ability, pooled over 41 studies by Freund and Kasten in 2012.
About 11%
Share of the variance in measured ability that a self-estimate accounts for at that correlation.
dz = 0.78
Size of the better-than-average effect across 291 samples and more than 950,000 people, Zell and colleagues 2020.
If you want the number rather than the guess, the honest route is a scored administration with a stated error term, and the honest caveat is that no online instrument, this one included, is a clinical or diagnostic tool. Before that, it helps to know what actually counts as a good score, because most people's intuition about the scale is as miscalibrated as their intuition about themselves.
The 1999 paper is four studies of Cornell undergraduates in four domains, and the effect it reported is smaller and more specific than the meme it produced. Kruger and Dunning published Unskilled and unaware of it in the Journal of Personality and Social Psychology, volume 77, number 6, pages 1121 to 1134. In each study participants took a test, then placed their ability and in most cases their test performance on a percentile scale running from 0 to 99.
Study 1 used humor. Sixty-five undergraduates rated 30 jokes, and their ratings were scored against a panel of professional comedians. The average participant placed their ability to recognize what is funny at the 66th percentile. The 16 participants in the bottom quartile had scored at the 12th percentile and placed themselves at the 58th, an overestimate of 46 percentile points. Even here, self-rated ability correlated .39 with the actual measure, which is the detail the meme drops.
Study 2 moved to logical reasoning with 45 participants. The 11 in the bottom quartile scored at the 12th percentile, placed their general reasoning ability at the 68th and their test score at the 62nd, and believed they had answered 14.2 problems correctly against an actual mean of 9.6. Study 3 used grammar with 84 participants, whose average self-placement was the 71st percentile for ability and the 68th for the test; its 17 bottom-quartile participants had scored at the 10th percentile and placed themselves at the 67th and the 61st. Study 4 returned to logic with 140 participants and added a training manipulation.
Study and domain
Participants
Bottom quartile actual
Bottom quartile self-placement
Study 1, humor
65
12th percentile
58th percentile for ability
Study 2, logical reasoning
45
12th percentile
68th for ability, 62nd for the test
Study 3, grammar
84
10th percentile
67th for ability, 61st for the test
Study 4, logic with training
140
13th percentile
55th for ability, 53rd for the test
What the paper did not claimThe idea that incompetent people believe they are experts is not in the data. In Study 4 the bottom quartile placed their ability at the 55th percentile and their test performance at the 53rd, and the authors report that neither figure was significantly greater than 50. The claim was miscalibration relative to a very low actual score, not delusions of expertise.
Two limits belong with these numbers. All four samples were Cornell undergraduates volunteering for course credit, a group with a compressed ability range, and the quartile cells contain 11 to 37 people. Percentile self-placement is also a comparative judgment that depends entirely on who the respondent pictures as the comparison group, which is why the difference between a score and a percentile matters more here than it looks. The same confusion runs in the other direction in public claims about other people: Muhammad Ali's documented 1964 result was a rank of about the 16th percentile on an armed forces qualifying test, and it reached print as an IQ of 78 with no stated assumption about which population or which distribution licensed the conversion, a case reconstructed under how a 16th percentile became an IQ of 78.
2 The Number That Settles the Practical Question
For anyone asking whether they can guess their own IQ, the quartile argument is beside the point, because the relevant statistic is the plain correlation between the guess and the score. Freund and Kasten located 41 published studies reporting 154 effect sizes and pooled them to r = .33. That estimate has since been stress tested rather than overturned. The 2024 multiverse meta-analysis by Patzl, Oberleiter and Pietschnig, Mirror, Mirror on the Wall in Journal of Intelligence volume 12, article 81, ran the pooling across nearly 5,000 defensible analytic specifications. Ninety-six percent produced a significant positive correlation, the specifications averaged r = .32, and the headline estimate was .30 with a confidence interval from .27 to .33.
A third, independent number points the same way. Reilly, Neumann and Andrews, publishing in Frontiers in Psychology in 2022 with a sample of 228 adults, found self-estimated intelligence correlating .30 with measured intelligence in their own data. Three separate research programmes using different samples and different instruments converge on a value near .3, which is unusual enough in psychology to be worth stating plainly.
What that value is not is a percentage of accuracy. A correlation of .33 does not mean self-estimates are 33 percent right. It means the two variables move together weakly, and both of them carry their own measurement error, so the true relationship between a person's belief and their standing is smeared by the imprecision of the instrument used to establish that standing. This is the same reason what makes a test accurate has to be argued from reliability evidence rather than asserted, and it is why the quality of the criterion test matters as much as the quality of the guess.
Two moderators from the Freund and Kasten abstract are worth carrying forward, because they are the only known levers. Validity improves when the self-estimate uses a relative scale with a clearly specified comparison group, and when the ability being estimated is numerical rather than general. Validity falls for dimensions people rarely think about, reasoning speed being the example the authors give. Patzl and colleagues add one refinement to the earlier work: correlations in student samples run significantly higher than in general population samples, which is a familiarity and range effect rather than a sign that students have better insight. None of this is enough to rescue the guess, and none of the instruments in a survey of online IQ tests can compensate for a weak starting signal.
3 What a Correlation of .33 Buys You in Points
Translate the correlation onto the scale people actually care about and the result is close to nothing. IQ composites are built with a mean of 100 and a standard deviation of 15. If a self-estimate correlates .33 with the score, then conditioning on the self-estimate leaves a residual standard deviation of 15 multiplied by the square root of 1 minus .33 squared, which is about 14.2 points. Knowing what somebody thinks their IQ is narrows the plausible range from 15 points to 14.2 points.
Expressed as an interval, that is worse still. A 95 percent window around an IQ predicted from a self-estimate alone spans roughly 55 points, wide enough to run from well below average to the edge of the range that gets called gifted. A prediction that cannot distinguish 85 from 135 is not a measurement, and no amount of introspective effort improves the arithmetic, because the ceiling is set by the correlation and not by the sincerity of the person doing the estimating.
About 14.2
Residual standard deviation in IQ points that remains after conditioning on a self-estimate, when r is .33.
About 55 points
Approximate width of a 95 percent interval around an IQ predicted from a self-estimate alone.
About 1.60
Standard error of measurement for the ACIS Full Scale IQ, derived from a composite omega of .9886.
The same arithmetic explains the pattern the Dunning-Kruger chart is famous for, without any psychology at all. When two variables correlate at .33, a person who is genuinely two standard deviations above the mean on one of them will, on average, sit only 0.66 standard deviations above the mean on the other. Put both on the familiar scale and that is a person near 130 who guesses near 110. Run it the other way and a person near 70 guesses near 90. The bottom overestimates and the top underestimates because the correlation is less than one, and that would remain true if every participant were perfectly honest and perfectly insightful about everything except the population distribution.
Against that, a scored administration reports an error term instead of a shrug. The ACIS Full Scale IQ has a composite omega of .9886 on the technical analysis set of 2,750 complete records, which yields a standard error of measurement of about 1.60 IQ points and a 95 percent interval roughly six points wide. Those figures come from internal consistency, so they understate total error, which also includes retest variation and testing conditions, and they were produced by unsupervised administration in a self-selected sample rather than a census panel. The comparison is still stark: an interval six points wide against one about 55 points wide. The short quiz sits between them rather than beside either: eighteen items of good quality reach a reliability near .76 and a band around 29 points, so three minutes of testing closes about half the distance from a guess to a measurement and leaves the rest of it standing, which is the arithmetic run instrument by instrument on how test length sets the error band. Once you have a number, how rare a score is and how the Full Scale IQ composite is assembled become answerable questions rather than speculation.
The crossing pair of lines that everyone recognizes is what you get whenever you sort people by one noisy measure and plot a second noisy measure against the sorted groups. The construction has three steps. Rank participants by their actual test score. Split them into quartiles. Then plot two things against those quartiles: the actual score, and the average self-estimate within each group.
The first of those lines is guaranteed to be steep, because the quartiles were defined by it. The second is guaranteed to be flatter, because a self-estimate correlates imperfectly with the score used to do the sorting, so group averages regress toward the overall mean. A flat line crossing a steep line produces exactly one picture: a large gap at the bottom pointing one way, a smaller gap at the top pointing the other, and a crossing point somewhere in the middle. The picture appears before anyone asks what the participants were thinking.
Kruger and Dunning saw the objection coming and addressed it in the paper. They wrote that the regression effect virtually guarantees the result, then argued that the miscalibration they observed was more psychological than artifactual, on the grounds that if regression alone were responsible the size of the miscalibration would be comparable at the top and the bottom, which it plainly was not. That argument is the load-bearing beam of the original interpretation, and it is the one the later literature attacked, because a second ingredient breaks the symmetry it assumes. If the whole distribution of self-estimates is shifted upward, the crossing point rises, the bottom gap widens and the top gap shrinks. Asymmetry no longer implies a metacognitive deficit.
What regression to the mean is notIt is not a claim that anybody's ability drifts toward the average over time, and it is not a property of IQ in particular. It is a fact about any two variables that correlate imperfectly, and it appears whenever a group is selected on one of them and then measured on the other. It is the reason how norms are built requires a representative sample rather than a convenience sample of extremes.
There is also a scale problem hiding in the percentile format. Participants who genuinely score at the second percentile cannot place themselves much lower than they actually are, but they have 97 points of room above. Participants near the ceiling have the opposite constraint. A bounded comparative scale therefore builds asymmetric error into the design, independent of insight. That constraint does not exist on the standard composite scale, where the population average of 100 sits in the middle of an unbounded normal distribution, which is one reason the two formats should not be discussed as though they were interchangeable.
5 The Simulations That Reproduced the Chart From Nothing
Nuhfer and colleagues took the argument out of the realm of debate by generating random numbers and showing which graphical conventions turn noise into a story. Their 2016 paper in Numeracy, Random Number Simulations Reveal How Random Noise Affects the Measurements and Graphical Portrayals of Self-Assessed Competency, volume 9, issue 1, article 4, states the premise directly: a self-assessment measure is a blend of an authentic signal and random disorder, and the choice of graph determines how much of the disorder gets read as signal.
The paper names the troublesome conventions explicitly. Scatterplots of self-assessment minus measured competence against measured competence. Column graphs of the same difference aggregated into quantiles. Line charts that display data aggregated as quantiles. Some histograms. The line chart of quantile averages is the Kruger and Dunning format. The conventions the authors found generate minimal artifacts are simple scatterplots of self-assessed against measured competence plotted by individual participant, and scatterplots of collective averages plotted item by item, the second of which attenuates noise and sharpens whatever signal is there.
The dataset behind both papers is substantial: 1,154 participants with paired measures of self-assessed and demonstrated competence in science literacy. The 2017 follow up, How Random Noise and a Graphical Convention Subverted Behavioral Scientists' Explanations of Self-Assessment Data, applies the better conventions to that dataset and reports a conclusion that contradicts the popular consensus: people's self-assessments of competence, in general, reflect a genuine competence that they can demonstrate. Same participants, same responses, different graph, opposite headline.
Read this claim narrowlyThe Nuhfer work is about graphical and numerical practice, not about intelligence. Its criterion measure is science literacy in an educational sample, not a standardised cognitive battery, and showing that a convention manufactures an artifact does not show that no one is ever overconfident. What it establishes is that the classic chart is not evidence for the classic explanation.
The general lesson transfers well beyond this dispute. A visual convention that is applied consistently across a literature can encode an assumption so thoroughly that nobody notices they are looking at the assumption rather than at the data. That is worth remembering when reading any strong claim about cognition drawn from group averages, and it belongs alongside the other items in myths about IQ testing that survived because the presentation was persuasive rather than because the evidence was.
6 Testing the Hypothesis Without Quartiles
If quartile charts are the problem, the fix is to test the underlying hypothesis with methods that never split the sample at all, and that is what Gignac and Zajenkowski did. Their 2020 paper in Intelligence, volume 80, article 101449, is titled The Dunning-Kruger effect is (mostly) a statistical artefact. They recruited 929 general community participants, not undergraduates, and collected a self-assessed intelligence measure alongside the Advanced Progressive Matrices.
The logic is neat. If accurate self-assessment genuinely depends on possessing the ability, then estimation error should be larger at the low end than at the high end, which is a prediction about the spread of residuals rather than about their average. That is testable directly with the Glejser test of heteroscedasticity. A second prediction, that the relationship between measured and self-assessed ability bends rather than running straight, is testable with nonlinear regression. Neither requires cutting anyone into a quartile.
Both tests came back negative. The authors report no statistically significant heteroscedasticity, contrary to the Dunning-Kruger hypothesis, and a relationship between objective and self-assessed intelligence that is essentially entirely linear. Their conclusion is carefully hedged: the phenomenon may hold for some specific skills, but the magnitude of the effect may be much smaller than previously reported. The instrument on the ability side was a matrix reasoning test, which is a strong marker of the general factor of intelligence but a narrow sample of it, a point that also applies when reading claims about Raven's 2 and its predecessors.
The artefact verdict has itself been challenged, and fairness requires saying so. Avram Hiller published a comment on Gignac and Zajenkowski in Intelligence volume 97 in 2023, arguing that the homoscedasticity finding is likely the result of a recoding choice: the authors mapped an ordinal comparative self-rating onto a linear IQ scale when a normal scale would probably have been more appropriate, and that transformation can flatten exactly the pattern the test was looking for. Hiller's broader warning is one this whole literature earns, that researchers studying self-assessed intelligence should watch for measurement problems introduced when an ordinal scale is pushed onto an interval one. So the position to hold is not that the effect has been disproved. It is that the strongest version of the original interpretation has lost its best evidence, while the argument about how to measure the thing at all remains open.
7 What Survives the Critique and What Does Not
The observed pattern is replicable, the popular explanation of it is not established, and there is a real metacognitive component that the critiques do not touch. McIntosh, Fowler, Lyu and Della Sala state the distinction cleanly in Wise up: clarifying the role of metacognition in the Dunning-Kruger effect, published in the Journal of Experimental Psychology: General in 2019, volume 148, pages 1882 to 1897. The pattern, they write, is not in doubt. The debate is about the correct explanation for it.
Their method is unusually direct. Participants pointed at a dot or recalled its position after a delay, with task skill measured in a first block and self-assessment measured in a second block, where they judged after every single trial whether they had hit the target. Metacognitive calibration and sensitivity did relate to task skill, but a path analysis showed their net contribution to the effect was weak, and the major driver was simply the level of task performance. In a second study the authors titrated difficulty so that everyone performed at equivalent levels of success, and the Dunning-Kruger pattern was eliminated. Their conclusion is that metacognitive differences can contribute to the effect but are neither necessary nor sufficient for it.
The strongest surviving evidence for a genuine metacognitive component comes from the original paper, in the part almost nobody cites. Study 4 randomly gave half the participants a training packet in logical reasoning. Bottom-quartile participants who received it went from correctly grading 3.5 of their 10 answers to correctly grading 9.3, which made them as accurate at monitoring their own performance as the participants who had originally scored in the top quartile. Their self-placement then fell from the 55th percentile to the 44th for ability, and from the 51st to the 32nd for the test. A randomly assigned intervention that moves a self-estimate cannot be regression to the mean, because regression does not respond to a manipulation.
Popular claim
Status
Basis
Incompetent people believe they are experts
Not what was reported
Study 4 bottom quartile placed itself at the 55th percentile, not significantly above the midpoint
The scissors chart demonstrates a metacognitive deficit
Not established
Nuhfer et al. 2016 and 2017 on graphical artifacts, McIntosh et al. 2019 on the weak net contribution
Error is larger at the low end of ability
Not supported in the one direct test
Gignac and Zajenkowski 2020 found no significant heteroscedasticity in 929 adults, though Hiller 2023 disputes the coding
Self-estimates track measured ability weakly
Established
Freund and Kasten 2012, Patzl et al. 2024, Reilly et al. 2022
Skill training improves self-assessment
Established experimentally
Kruger and Dunning 1999, Study 4 training condition
High performers underestimate their standing
Reported in every study
Kruger and Dunning 1999, Studies 1 to 4
None of this depends on a person's standing being summarised by one number. A test that reports the six broad ability domains separately gives a reader something a composite cannot, and knowing the kinds of subtest inside a battery is what makes a profile interpretable rather than decorative.
8 The Better Than Average Effect Does the Other Half
The second ingredient in the classic chart is the most robust finding in social psychology that almost nobody names when they discuss Dunning-Kruger. Zell, Strickhouser, Sedikides and Alicke published the first quantitative synthesis of the better-than-average effect in Psychological Bulletin in 2020, volume 146, pages 118 to 149, pooling 124 articles across 291 independent samples and more than 950,000 participants. The effect size was dz = 0.78 with a confidence interval from .71 to .84, and the authors reported little evidence of publication bias overall.
The mechanism it supplies is exactly what the regression argument needs. If most people place themselves above the midpoint before any question about their actual standing arises, then the entire distribution of self-estimates sits high. Combine a shifted distribution with a flat regression line and the crossing point moves upward, which widens the apparent overestimation at the bottom and shrinks the apparent underestimation at the top. That is the asymmetry Kruger and Dunning treated as evidence that something psychological beyond regression was happening. It is psychological, but it is a different psychology from the one they proposed, and it is not concentrated in low performers.
Two moderators from the same meta-analysis complicate the picture in useful ways. The effect was larger for personality traits than for abilities, which is a genuine limitation on using it to explain an ability chart, and it was larger in European American samples than in East Asian samples, particularly for individualistic traits. A phenomenon that varies that much by culture is not a fixed feature of human cognition, and a self-estimate collected in one country cannot be read as though it meant the same thing collected in another.
The same paper reports that the better-than-average effect correlates .34 with self-esteem and .33 with life satisfaction. Those numbers are the same size as the correlation between self-estimated and measured intelligence, which is a blunt way of saying that a person's guess about their own IQ is roughly as informative about how they feel as it is about how they score. Anyone tempted to read a confident self-estimate as a personality reading rather than an ability reading will find the evidence on how personality relates to measured ability more useful than the guess itself.
9 The Half of the Finding Everyone Forgets
In every one of the four 1999 studies, participants in the top quartile placed themselves below where they actually stood, and that half of the result never made it into the meme. In Study 1 the top quartile underestimated their ability relative to their peers, with a paired t of minus 2.20 at p below .05. In Study 2 they had scored at the 86th percentile and placed their reasoning ability at the 74th and their test score at the 68th. In Study 3 the 19 top-quartile participants had scored at the 89th percentile and placed their grammar ability at the 72nd and their test performance at the 70th. In Study 4 they had scored at the 90th percentile and placed their ability at the 76th and their test performance at the 79th.
Part of that is the arithmetic already described: with an imperfect correlation, high scorers regress downward in their estimates just as low scorers regress upward. But the original paper offered a psychological account too, and it holds up better than the dual burden account it accompanied. High performers find the task easy, assume it was easy for everyone, and therefore place themselves closer to the middle than they belong. The error is about the peer group, not about themselves.
This has a practical consequence that the popular framing inverts. The person who is most confident that they are unexceptional may be sitting well above the range they imagine, and the feeling of ordinariness is not evidence. It is particularly unreliable inside a selective environment, where the available comparison group is a filtered slice of the population rather than the population. Someone surrounded by graduate colleagues will produce a lower self-estimate than the identical person in a general workplace, without either estimate being dishonest.
That reference group problem is why phrases like the gifted range only mean something when the reference frame is stated, and why record claims at the top of the scale collapse when you ask what norms they were scored against. It is also why the behavioural cues collected under signs people read as intelligence are a poor substitute for a score, and why an organisation such as Mensa admission testing uses a supervised administration and a fixed cutoff rather than asking applicants what they think.
10 The Gender Gap Is in the Estimate, Not the Score
The best replicated moderator of self-estimated intelligence is sex, and it appears in the estimate rather than in the measurement. Szymanowicz and Furnham ran four meta-analyses on this question in Learning and Individual Differences in 2011, volume 21, pages 493 to 504. As reported by Reilly and colleagues in 2022, the weighted mean effect sizes favoured males at d = .37 for general intelligence, d = .44 for mathematical and logical ability, and d = .43 for spatial ability, with a much smaller d = .07 for verbal ability.
A more recent primary study puts the same pattern on the familiar scale. Reilly, Neumann and Andrews published Gender Differences in Self-Estimated Intelligence: Exploring the Male Hubris, Female Humility Problem in Frontiers in Psychology in 2022, with 228 participants, 103 male and 125 female. Mean self-estimated IQ was 112.12 with a standard deviation of 9.20 for men and 103.66 with a standard deviation of 10.88 for women, a gap of about 8.5 points and an effect size of d = 0.74. In the same sample, self-estimated intelligence correlated .30 with measured intelligence for everybody.
Two things follow, and the order matters. The gap is in what people say, and the whole point of the .30 correlation is that what people say is a poor guide to what they score. The authors describe the effect as multifactorial rather than a simple product of biological sex, involving masculine gender role traits and self-esteem as contributors, and they explicitly reject social desirability as a sufficient explanation. Their own framing is that the label describes a pattern of estimates, not a pattern of abilities.
The boundary this page will not crossNothing in this section is a claim about mean differences in measured cognitive ability between groups. The evidence cited here is evidence about self-estimates and nothing else. Group statistics of any kind are also never a description of an individual, and readers interested in how instruments are examined for differential functioning should start with the questions raised about test bias rather than with self-report data.
One practical implication is uncomfortable and worth stating. If two people of identical measured ability produce self-estimates 8.5 points apart on average, then any process that relies on self-assessed ability, from a course placement conversation to a candidate deciding whether to apply, will systematically misroute one of them. This is the strongest non-commercial argument for measuring rather than asking, and it applies with equal force to the informal version, where someone decides they are or are not the sort of person who takes an adult IQ test at all.
11 What Self Estimates Actually Track
If a self-estimate shares only about a tenth of its variance with measured ability, the other nine tenths are made of something, and the literature is fairly clear about what. Freund and Kasten give the answer in their own abstract: self-estimates provide diagnostic information about a person's self-concept, which they suggest is useful in career counselling and educational settings. That is a real use, and it is not the use most readers came looking for.
The correlations line up with that reading. The better-than-average effect correlates .34 with self-esteem and .33 with life satisfaction in the Zell meta-analysis, essentially the same magnitude as the correlation between self-estimated and measured intelligence. Reilly and colleagues found masculine gender role traits and self-esteem contributing to self-estimated intelligence alongside sex. A number that predicts your mood about as well as it predicts your score is not primarily a measurement of your score.
The cue problem explains the rest. People assemble a self-estimate out of school grades, highest qualification, job title, whether they were the fast one in a particular classroom, and comments made by a handful of people who know them. Every one of those cues is genuinely correlated with measured ability, which is what makes the estimate better than chance. Every one is also confounded with opportunity, health, language background, effort and luck, which is what keeps it at .3. Building a guess from cues like average scores by education level produces a circular estimate, since educational attainment is partly a consequence of the very thing being estimated. Each of those cues is stamped as well with the age at which it was collected, while a score is a rank inside an age band: rank order correlates .67 between ages 11 and 70 even as raw performance on speeded and unfamiliar material falls for decades across the same span, so a fifty year old reasoning from what came easily at twenty is answering with a memory of a different comparison group, a split set out under what ageing moves and what it leaves in place.
It is also worth separating the constructs people quietly substitute. The evidence on the link with academic achievement and the link with job performance describes group level associations with wide individual scatter, so working backwards from an outcome to a score is unreliable in both directions. And a good deal of what people mean by feeling smart or feeling slow belongs to other constructs entirely, which is why emotional intelligence compared with IQ keeps coming up whenever a self-estimate and a measured score disagree. The labels get borrowed as loosely as the numbers do, since the 140 genius threshold that anchors most private talk about high ability comes from a single suggested table in Terman's 1916 manual, a row he had dropped from his own 1937 revision, and the trail behind that threshold ends there rather than in any instrument now in print. A self-estimate is a summary of a self-concept. It behaves like one.
12 The Conditions That Make a Self Estimate Less Bad
The meta-analytic moderators identify real conditions under which self-estimates improve, and they are specific enough to act on, though not specific enough to rescue the exercise. Freund and Kasten found validity enhanced when relative scales with clearly specified comparison groups are used, and when the ability in question is numerical rather than general. They found it reduced for dimensions people rarely consider, with reasoning speed as their example.
Translated into practice, that means the question matters. Asking somebody how smart they are invites a self-concept report. Asking them to place their arithmetic ability against a named and imaginable group, such as people who finished secondary school in their own country, invites something closer to a judgment. Narrow beats broad, named comparison groups beat vague ones, and a domain the person has had repeated concrete feedback in beats one they have never been scored on.
The improvement is still not enough. Suppose the favourable conditions lift the correlation from .33 to .45, which would be at the optimistic end of what the moderator analyses suggest. The residual standard deviation falls from about 14.2 points to about 13.4. A 95 percent interval narrows from roughly 55 points to roughly 52. Nothing in that range of values turns a self-estimate into a basis for a decision, which is why the entire discussion ends where it does.
There is one form of self-knowledge that holds up better, and it is relative rather than absolute. People are markedly better at ranking their own abilities against each other than at placing themselves in a population. Knowing that your verbal reasoning outruns your spatial reasoning is a within-person comparison, and it does not require an accurate model of the population distribution to be useful. That is precisely the information a domain profile carries and a single composite hides. A battery organised on the CHC model of cognitive abilities reports VCI, FRI, QRI, VSI, WMI and PSI separately, so that Gs appears on its own line with processing speed treated as a domain rather than absorbed into a general number, and the split between fluid and crystallized ability becomes visible instead of averaged away. Dispersion across those indices is normal, not a defect, and reading it well is a different skill from critical thinking as a separate capacity, which these tests do not sample at all.
13 The Arithmetic Case for Measuring, and Its Limits
If the correlation between a self-estimate and a measurement is about .3, then the only way to find out is to measure, and that sentence is arithmetic rather than a sales argument. It carries no implication that the number is important, that it should change any decision, or that a particular instrument is the right one. It says only that a question about where you stand cannot be answered by consulting your impression of where you stand.
Property
A self-estimate
A scored administration
Reference group
Whoever the person happens to picture
A stated reference frame, here 3,243 English speaking records aged 16 to 90
Reliability evidence
None exists
Composite omega of .9886 for the Full Scale IQ on 2,750 complete records
Stated error
Not defined
Standard error of measurement of about 1.60 IQ points
Structural evidence
Not applicable
Higher order model fit of CFI .9761, TLI .9726, RMSEA .0406, SRMR .0217, chi-square 916.703 on 166 degrees of freedom
Detail available
One impression
20 subtests across six domains, scaled scores at mean 10 and SD 3, composites at mean 100 and SD 15
Known limitations
Correlates about .3 with the thing it claims to describe
Self-selected sample, unsupervised administration, a modelled adult reference frame rather than a census sample
Those limitations are not fine print. The reference frame is built from people who chose to take a test online, which is not a random slice of the adult population, and administration is unsupervised, so identity, environment and effort are not controlled the way they are when a psychologist sits in the room. The g loading of .958 for the composite and the fit statistics above describe the internal structure of the instrument in that sample; they do not convert an unproctored session into a supervised professional assessment. Anyone who needs a score for a clinical question, an educational accommodation, an employment decision or admission to a high IQ society should read where a supervised test can be taken and go there instead. Comparison with a proctored battery such as the WAIS-5 is a comparison of administration conditions before it is a comparison of content.
The honest close is that a measured score answers a narrower question than the one most people bring to it. It will not tell you whether you are good at your job, whether your reasoning about a specific argument is sound, or whether you are competent in a domain you have never studied. It gives a position on a defined scale with a defined error term, which is exactly the thing a self-estimate cannot give and exactly the thing the Dunning-Kruger meme promises to diagnose from the outside. Any unsupervised online administration, and any use of the same instrument through the research and screening workspace or by someone administering the test to other people, inherits both the arithmetic and the limits.
That boundary is not a house style. The Standards for Educational and Psychological Testing (2014), issued jointly by the American Educational Research Association, the American Psychological Association and the National Council on Measurement in Education, make validity a property of an interpretation for a specified use rather than a property a test possesses in general, and they require that reliability evidence and an error term accompany a reported score. The APA testing standards overview restates the same obligation for practitioners. Read through that framework, a self-estimate is not a weak test score, it is not a test score at all: it has no norm sample, no reliability coefficient and no stated error, and nothing in the Standards would license using one to place a person anywhere.
What is the Dunning-Kruger effect in one sentence?
It is the reported finding that people who perform worst on a task overestimate their standing by the largest margin, while the best performers slightly underestimate theirs. The pattern replicates. The explanation originally attached to it is what later work has disputed.
How accurate are people at estimating their own IQ?
Poorly, and roughly equally poorly across the range. Pooled correlations between self-estimates and measured ability sit near .3 in every large synthesis published so far, which accounts for about a tenth of the variation between people.
Is the Dunning-Kruger effect real or a statistical artefact?
Partly both. The graph is largely reproducible from regression to the mean plus a general tendency to rate oneself above the midpoint, but a genuine metacognitive contribution remains, demonstrated experimentally rather than by the chart.
Do smart people underestimate their own intelligence?
On average, yes, and that half of the original finding is rarely quoted. Top-quartile participants placed themselves below their actual standing in all four of the 1999 studies, by 11 to 19 percentile points wherever the paper prints the figures.
What does a correlation of .33 mean in IQ points?
It reduces the uncertainty about someone's score from a standard deviation of 15 points to about 14.2. A 95 percent interval predicted from a self-estimate alone spans roughly 55 points, which is too wide to separate below average from gifted.
Did Kruger and Dunning say incompetent people think they are geniuses?
No. In their fourth study the lowest-scoring quarter placed themselves at the 55th percentile, and the authors report that figure was not significantly above the midpoint. The claim was miscalibration against a very low score, not delusion.
Who took part in the original studies?
Cornell University undergraduates earning course credit, in samples of 65, 45, 84 and 140. That is a narrow and academically selected group, which compresses the ability range and limits how far the findings generalise to adults at large.
What is regression to the mean and why does it matter here?
When two variables correlate imperfectly, extreme scores on one are matched by less extreme scores on the other. Sorting people by test score and plotting their estimates therefore produces a flatter line by construction, before any psychology is involved.
What did the random number simulations show?
Nuhfer and colleagues showed that specific graphical conventions, particularly line charts of quantile-aggregated differences, generate artifacts that invite misreading. Plotting individual participants instead, on the same data, supports a very different conclusion about self-assessment.
What did Gignac and Zajenkowski find?
In 929 community adults they tested the hypothesis without quartile splits, using a heteroscedasticity test and nonlinear regression. Estimation error did not grow at the low end, and the relationship with measured ability was essentially linear throughout.
Has the artefact claim itself been challenged?
Yes. Avram Hiller argued in 2023 that the finding depends on how an ordinal comparative self-rating was transformed onto a linear scale, and that a different transformation could restore the pattern. The methodological dispute has not been settled.
Does training improve self-assessment?
It did in a controlled test. Bottom-quartile participants given a logical reasoning packet went from grading 3.5 of their own answers correctly to 9.3, and lowered their self-placement accordingly. A random assignment effect cannot be regression to the mean.
Why do men give higher self-estimates than women?
The reasons look multifactorial. Researchers point to masculine gender role traits and self-esteem alongside sex itself, and they reject social desirability as a sufficient account. The gap is well replicated across countries, ages and ethnic groups.
Does the gender gap in self-estimates reflect a gap in measured scores?
No, and nothing on this page claims otherwise. The evidence cited concerns what people say about themselves. Since self-estimates track measured ability at only about .3, a difference in estimates carries almost no information about scores.
What is the better-than-average effect?
The tendency to rate oneself as superior to the average peer. Synthesised across 291 samples and over 950,000 people, its size is dz = 0.78. It is larger for personality traits than abilities and varies substantially by culture.
If I feel average, does that mean I am average?
It is weak evidence at best. Feeling ordinary is exactly what the arithmetic predicts for someone well above the mean, and selective environments push the impression lower still by supplying an unrepresentative comparison group.
Does a high self-estimate mean I am overconfident?
Not on its own. Most people place themselves above the midpoint, so a high estimate is normal rather than diagnostic. It becomes informative only when set beside a score, and even then it says more about self-concept than about calibration.
Are self-estimates useless for anything?
No. Freund and Kasten describe them as diagnostic information about self-concept, with applications in career counselling and education. They measure something real. That something is how a person sees themselves, not where they sit on a scale.
Does knowing a subject well make my self-estimate better?
Somewhat. Estimates improve for narrow, concrete abilities such as numerical skill and when the comparison group is clearly specified. The gain is real but small, and it never approaches the precision that any actual decision would require.
Can an unsupervised online test tell me my IQ?
It can produce a scored position with a stated error term, which a guess cannot. It cannot control identity, environment or effort, it is not a clinical or diagnostic instrument, and its reference frame is self-selected rather than a census sample.
What should I do if my measured score is far from my guess?
Treat the gap as expected rather than alarming, since the two things correlate weakly by design. Read the index profile rather than the composite alone, and check the error interval before treating any single point value as fixed.
Take the assessment
You get a profile, not a number
ACIS measures six CHC domains across 20 subtests and reports each one with its own normed score and confidence interval, so you can see where you are strong and where you are not.