The claim that funny people are smarter is one of the few popular intelligence claims with real data behind it. The data is narrower than the headline: it is about producing humor rather than enjoying it, it runs mostly through verbal ability, and the funniness in every study was scored by strangers rather than by the person who wrote the joke.
Across seven studies of 1,133 adults, how funny people judged their own attempts correlated .13 with how funny independent raters judged them.
0 Quick Answer
Being funny does track measured ability, but only when funniness is measured by having people write jokes that strangers then rate, and the association runs through verbal knowledge rather than through reasoning with unfamiliar material. The cleanest decomposition is Christensen, Silvia, Nusbaum, and Beaty, reported in Psychology of Aesthetics, Creativity, and the Arts in 2018 on 270 adults. Humor production ability correlated .49 with crystallized ability, .38 with broad retrieval ability, and .22 with fluid reasoning. A general factor estimated from all three predicted humor production at .51.
That is a substantial relationship by the standards of individual differences research. It is also a relationship with a specific shape that popular coverage flattens. Producing humor is not the same skill as appreciating it, and the evidence for the two is not equally strong. The correlation sits with verbal ability more than with fluid reasoning, and in a bifactor model the fluid effect disappears entirely once general ability is accounted for. And humor production is measured with cartoon caption tasks scored by trained raters, never by asking people whether they are funny, because self-ratings barely track judged performance.
This page is about the research on humor and measured ability. Where humor sits among the behavioral markers people look for is covered separately in the list of signs of high intelligence, which owns that framing. What follows here is the studies, their samples, their effect sizes, and the two places where the popular version of this finding goes wrong.
.49 versus .22
Latent correlations of humor production ability with crystallized ability and with fluid reasoning, Christensen and colleagues, 2018, n = 270.
.38
Correlation between vocabulary score and rated caption funniness across 400 students, Greengross and Miller, Intelligence, 2011.
.13
Within-person correlation between self ratings and judge ratings of the same humor attempts, Silvia and colleagues, 2021, across seven studies and 1,133 adults.
Every result on this page rests on the same procedure, and understanding the procedure explains most of what the results can and cannot support. Participants are given a stimulus, usually a cartoon with its caption removed, and told to write something funny. Their attempts are then rated for funniness by judges who do not know who wrote them and who never see the writer's test scores. The rated funniness is the dependent variable. Nobody is ever asked whether they have a good sense of humor.
The details vary and the variation matters. Greengross and Miller used captions for three cartoons rated by six judges, and reported that interrater agreement averaged 0.72 across the three. Christensen and colleagues used three different task types, cartoon captions, joke stems, and satirical definitions, and scored them with a many facet Rasch model that adjusts each person's estimate for how lenient or severe the raters who happened to see it were. Their person reliability estimates were .68 for definitions, .63 for jokes, and .45 for cartoon captions, which is a direct warning that the single most popular task in this literature is also the noisiest one.
Howrigan and MacDonald went further from the cartoon convention, using six character profiles, three sarcastic email responses, and two pages of humorous drawings. Their eleven item composite reached a Cronbach's alpha of .74 with a mean interrater correlation of .42, while the individual items averaged .58 and .28. That gap is the whole reason these studies aggregate across many attempts: one joke tells you almost nothing about a person, and eleven of them tell you something.
What the measure cannot reachRated funniness on a written task in a laboratory is not the same construct as being funny in a conversation. Timing, delivery, reading a room, and knowing when not to make the joke are all absent from every study cited here. What is being measured is the ability to generate an amusing idea in writing on demand, which overlaps with everyday humor without being identical to it. No author in this literature claims otherwise.
The consequence is that a low score on one of these tasks is weak evidence about a person and a group difference across hundreds of them is much stronger evidence about a population. That asymmetry runs through every outcome finding, and the general logic of it is set out on the page on measured ability and life outcomes.
2 Producing Humor and Enjoying It Are Different Findings
The single most common error in coverage of this topic is treating a study of humor production as though it were a study of humor appreciation, or the reverse. They are different tasks with different demands and different evidence behind them, and the strong result belongs to production.
Production means generating something funny that did not exist a moment ago. It requires retrieving relevant material from memory, holding several possible directions in mind, selecting the one with the sharpest incongruity, and expressing it compactly. Appreciation means recognizing that something already constructed is funny. The construction work has been done for you; what remains is detecting the incongruity and resolving it.
Every effect size on this page above .20 comes from a production study. The appreciation literature is smaller, uses ratings of existing material, and reports weaker and less consistent relationships with measured ability, in part because appreciation ratings are compressed. Most people find most professionally written jokes at least somewhat funny, so there is less variance to explain. Production tasks spread people out because most attempts fail.
Production tasks give a stimulus and require an original response, which is then rated by strangers. This is where the .29 to .51 range of effects comes from.
Appreciation tasks present finished material and record ratings of funniness, difficulty, or comprehension. This is what the black humor research measured.
Self-report humor scales ask people about their own humor styles or abilities. These measure beliefs about oneself, which is a different variable again.
Keeping the three apart resolves most of the apparent contradictions in popular articles on this topic. A study finding that people who enjoy dark humor score higher and a study finding that people who write funnier captions score higher are not two measurements of the same thing, and neither one licenses the other's headline. The same distinction between generating and recognizing runs through the adjacent literature on measured ability and creativity, where divergent production and recognition tasks also behave differently.
3 The Study That Set the Reference Point
Greengross and Miller established the association in a sample large enough to split by sex and still have power in each half. Their paper appeared in Intelligence in 2011, volume 39, pages 188 to 192. Four hundred University of New Mexico students took part, 200 men and 200 women. Abstract reasoning was measured with Raven's Advanced Progressive Matrices, verbal ability with the vocabulary subtest of the Multidimensional Aptitude Battery, and humor production with the rated funniness of captions written for three cartoons.
The headline correlations were .38 between vocabulary score and humor ability, and .25 between Raven's score and humor ability, both at p below .001. The authors tested whether the difference between those two was real using a Fisher r to z transformation and reported z = 2.1, significant at p below .05. That single comparison is the origin of the claim that the humor link is verbal, and it is worth noting that the authors ran the test rather than eyeballing the gap.
Split by sex, the pattern held in both halves. For men, humor ability correlated .44 with vocabulary and .27 with Raven's. For women, the same correlations were .31 and .24. The structural equation models produced latent effects of general intelligence on humor ability of .67 for men and .51 for women, figures later described by Christensen and colleagues as among the large effects in this literature.
The sex difference in humor production itself was smaller than the intelligence association. Men averaged 0.09 on the standardized humor measure with a standard deviation of 0.48, women averaged minus 0.09 with a standard deviation of 0.49, giving t = 3.77 and a Cohen's d of 0.38. That is a real difference in the sample and a modest one, with the two distributions overlapping across most of their range.
The limits the authors named themselvesGreengross and Miller wrote that the sample was drawn from one university, that future work should use larger community samples across cultures and a wider age range, and that a larger number of intelligence measures would be needed before a proper general factor could be extracted. The .38 and .25 figures therefore describe undergraduates at a single institution, not adults in general.
4 The Study That Took the Association Apart
The most informative paper in this literature is the one that stopped asking whether intelligence predicts humor and started asking which ability does. Christensen, Silvia, Nusbaum, and Beaty ran 270 young adults at the University of North Carolina at Greensboro through three humor tasks and three families of cognitive tasks, then modeled the result inside the CHC framework. Mean age was 19.08 with a standard deviation of 3.14, and 86 percent of the sample was women.
Fluid reasoning was measured with paper folding, series completion, letter sets, and number series. Crystallized ability was measured with two vocabulary tests. The addition that made the paper matter was broad retrieval ability, measured with five fluency tasks in which participants generate as many instances of a category as they can. That factor had never been examined in humor research before.
The latent correlations came out as follows. Crystallized ability correlated .49 with humor production, with a confidence interval from .29 to .70. Broad retrieval correlated .38, interval .22 to .54. Fluid reasoning correlated .22, interval .04 to .39, the only one of the three whose interval came close to zero. A confirmatory model fit the data well, with a chi-square of 112.768 on 72 degrees of freedom, a CFI of .937, an RMSEA of .046, and an SRMR of .050.
.51
Standardized effect of a higher order general factor on humor production ability, with a confidence interval from .32 to .70.
.20
The unique effect of broad retrieval ability that survived after general ability was modeled, interval .02 to .38.
.09
The unique effect of fluid reasoning after general ability was modeled, p = .398, which is to say no unique effect at all.
That last row is the finding that changes how the whole literature should be read. In the bifactor model, the general factor predicted humor production at .47, broad retrieval kept a specific effect of .20, and fluid reasoning's specific effect collapsed to .09 and lost significance. The authors state the implication directly: the variance in the fluid tasks that predicts humor ability is the variance due to the general factor, not the variance unique to fluid reasoning. Crystallized ability, meanwhile, was absorbed into the general factor entirely, because the general factor explained the vocabulary tasks so thoroughly that no separate crystallized factor remained.
5 Every Study Side by Side
Laid out together, the studies agree on the direction and disagree on the size in ways that track their measures. The pattern across four decades is a correlation somewhere between .22 and .51 depending on which ability is measured, how many humor attempts are aggregated, and whether the estimate is a raw correlation or a latent one corrected for measurement error.
Study
Sample
Humor measure
Ability measure
Reported association
Masten, 1986
Children aged 10 to 14
Rater-judged humor production
IQ and academic achievement
r of .50 to .53, as reported by Howrigan and MacDonald
Feingold and Mazzella, Personality and Individual Differences, 1991
College age sample
Rater-judged verbal humor production
Measured verbal intelligence
r = .40, as reported by Howrigan and MacDonald
Howrigan and MacDonald, Evolutionary Psychology, 2008
185 students, 115 women and 70 men
11 items: character profiles, email replies, drawings
12-item short form of Raven's APM
r = .29 raw; beta of .20 in a ten predictor regression
Greengross and Miller, Intelligence, 2011
400 students, 200 men and 200 women
Captions for three cartoons, six judges
MAB vocabulary and Raven's APM
r = .38 verbal, r = .25 abstract; latent .67 and .51 by sex
Kellner and Benedek, Psychology of Aesthetics, Creativity, and the Arts, 2017
151 participants
Punch lines for six caption-removed cartoons
Crystallized and fluid measures
r = .37 crystallized, r = .13 fluid, as reported by Christensen and colleagues
Christensen, Silvia, Nusbaum, and Beaty, 2018
270 young adults
Captions, joke stems, definitions, Rasch scored
Gf, Gc, and Gr latent factors
Gc .49, Gr .38, Gf .22; general factor .51
Willinger and colleagues, Cognitive Processing, 2017
156 adults, 76 women and 80 men
Ratings of 12 black humour cartoons, not production
Vocabulary test and a number connection test
Group means differing by 9 to 20 points across three clusters
Two features of the table deserve attention. The first is that the two studies using a single narrow humor task, Howrigan and MacDonald with Raven's alone and Kellner and Benedek with a fluid measure, produce the smallest correlations, which is what measurement error does to an estimate. The second is that the last row is not measuring the same thing as any row above it, and that is the subject of its own section below.
Howrigan and MacDonald's regression is worth extracting because it addresses the obvious confound. They entered ten predictors at once: general intelligence, all five Big Five traits, age, height, weight, and semesters in college. Two survived. General intelligence had a standardized beta of .20 at p = .010 and extraversion had .16 at p = .037. Openness, conscientiousness, agreeableness, and neuroticism did not predict rated funniness once intelligence was in the model. The relationship between measured ability and trait structure generally is covered on the personality page.
6 Why the Association Is Verbal Rather Than Fluid
The tilt toward verbal ability is not a quirk of the samples; it follows from what a joke requires. Writing a caption means retrieving candidate meanings for the objects in the picture, finding a second frame in which those meanings mean something else, and compressing the collision into one line. Every step of that is retrieval and manipulation of stored knowledge.
Consider what the fluid tasks in these studies actually demand. Paper folding, letter sets, series completion, and matrix reasoning all present novel material with no prior knowledge required, and reward finding the rule. That is a real ability and it is the backbone of most composites, but it is not the ability that supplies the second meaning of a word. Vocabulary tests, by contrast, are direct samples of exactly the store a joke draws on. When Greengross and Miller found .38 for vocabulary against .25 for Raven's and confirmed the gap at z = 2.1, they were finding a difference their materials predicted.
Christensen and colleagues added the piece that makes the mechanism concrete. Broad retrieval ability, measured with tasks such as naming as many animals or as many things that are hot as possible in a fixed time, kept a unique effect of .20 on humor production even after general ability was modeled, while fluid reasoning kept nothing. Retrieval fluency is the ability to search semantic memory quickly and strategically, which is the operation a punchline needs most and the one a matrix problem needs least.
The distinction between reasoning with unfamiliar material and drawing on accumulated knowledge is the oldest division in the ability literature, and it is set out on the fluid versus crystallized page. The specific tasks that sample the verbal side, and how they differ from each other, are described on the verbal ability page and in the subtest descriptions for Vocabulary, Similarities, and Synonyms.
An important gap in what ACIS measuresACIS has no retrieval fluency subtest. Its 20 subtests cover verbal comprehension, fluid reasoning, quantitative reasoning, visual spatial processing, working memory, and processing speed, but nothing in the battery asks a person to generate as many instances of a category as they can. The .20 unique effect that Christensen and colleagues attributed to broad retrieval therefore has no direct counterpart in an ACIS profile, and this page does not claim one.
7 The Dark Humor Headline and What the Study Found
The most widely circulated claim in this area comes from one study of 156 people that measured comprehension of cartoons, not the ability to be funny, and the press coverage overstated it in three separate ways. Willinger and colleagues published Cognitive and emotional demands of black humour processing in Cognitive Processing in 2017, volume 18, pages 159 to 167.
Here is what the study did. One hundred and fifty-six adults, 76 women and 80 men, with a mean age of 33.4 and a standard deviation of 11.9, rated 12 black humour cartoons drawn from a collection by Uli Stein. Six of the cartoons dealt with death, three with physical handicap, two with disease, and one with medical treatment. Participants rated each cartoon on difficulty, fit, preference, and comprehension. Verbal ability was measured with a German vocabulary test and what the paper calls nonverbal IQ was measured with the Zahlen-Verbindungs-Test, a number connection task.
A cluster analysis split the sample into three groups. Group one, 41 people, had a verbal mean of 101.0 and a nonverbal mean of 97.8. Group two, 50 people, had 96.8 and 102.8. Group three, 65 people, the group with the highest preference for and comprehension of the cartoons, had 109.7 and 118.1. The differences were significant, F(2,153) = 112.6 for the nonverbal measure and F(2,153) = 19.9 for the verbal one, both at p at or below .0001. Group three also contained significantly more higher educated participants at p = .014.
Now the three overstatements. First, the study is about comprehension and preference, not production; nobody in it wrote a joke. Second, the group means are not high in absolute terms. A group mean of 118 on a speeded number connection task is High Average, not gifted, and it is a group average rather than a description of any individual. Third, the task labeled nonverbal intelligence is a trail making test in which the participant connects numbers in sequence as fast as possible, developed by Oswald and Roth within the mental speed tradition. It is much closer to a measure of processing speed than to a measure of reasoning, so reporting the result as a difference in nonverbal intelligence claims more construct coverage than the instrument delivers.
The variable that moved most is not intelligenceMood disturbance separated the groups more sharply than either ability measure. Group two averaged 19.5 with a standard deviation of 11.5, more than double group one at 9.2 and group three at 8.6. Aggressiveness toward others followed the same order, 14.5 against 11.2 and 9.4. The group that liked black humour least was distinguished mainly by being in a worse mood, and a headline about intelligence skips the largest effect in the paper.
8 People Are Poor Judges of Their Own Jokes
Asking someone whether they are funny produces a number that barely relates to whether other people find them funny, which is why no serious study in this literature uses self-report as the outcome. The best evidence for this is Silvia, Greengross, Cotter, Christensen, and Gredlein, published in Journal of Research in Personality in 2021, volume 92, article 104089.
Across seven studies with 1,133 adults, participants created funny ideas, rated the funniness of their own responses, and had those same responses independently rated by judges. The within-person correlation between the self ratings and the judge ratings was .13, with a confidence interval from .07 to .19 across all seven studies. The authors describe this as small but significant, meaning people have some insight into which of their own attempts landed, and not much.
The paper also found that personality predicts self-rating rather than performance. Extraversion correlated .12 with rating one's own responses as funnier, interval .07 to .18. Openness correlated .09, interval .03 to .15. Women rated their own responses as less funny than men rated theirs, with a d of minus 0.28, interval minus 0.37 to minus 0.19. None of those associations is with judged funniness. They are associations with confidence.
One further result cuts against the usual story. The well known finding that most people rate themselves as funnier than average did not appear when people rated specific attempts. Asked about a particular joke they had just written, participants were relatively modest and self-critical. The inflation lives in the global self-judgment, not in the local one.
That pattern generalizes well beyond humor. Self estimates of ability in general correlate loosely with measured ability, which is the reason a self-assessment is not a substitute for an administration. The evidence on that gap is collected on the page on estimating your own ability.
9 What the Comedians Data Adds
One study went to the people who do this for a living, and the result is consistent with the student data while being much smaller and much harder to interpret. Greengross, Martin, and Miller published their comparison of professional stand-up comedians with college students in Psychology of Aesthetics, Creativity, and the Arts in 2012, volume 6, pages 74 to 82.
Thirty-one professional comedians were compared with 400 college students on the Big Five, the Humor Styles Questionnaire, a humor production task, and a measure of verbal intelligence. The comedians scored higher than the students on verbal intelligence, on humor production ability, and on each of the four humor styles the questionnaire measures. Within the comedian group, openness, agreeableness, and extraversion correlated positively with affiliative humor, and intelligence correlated negatively with self-defeating humor.
Take the sample size seriously before taking the result seriously. Thirty-one people is enough to detect a large group difference and not enough to estimate a correlation within the group with any precision. A within-comedian correlation from 31 people has a confidence interval wide enough to include almost any moderate value, so the personality findings inside that group should be read as suggestive rather than established.
There is also a selection issue that no design can remove. Professional comedians are people who tried comedy, were good enough to keep going, and stayed in the profession long enough to be recruited for a study. Any comparison between them and undergraduates confounds ability with survivorship, motivation, and years of deliberate practice. The finding that they score higher on a verbal measure is compatible with the ability being a cause, with the practice being a cause, and with both.
The honest summary is that the comedian data is consistent with the student data and adds little independent weight. What it does add is a reminder that the constructs in this literature are separable: humor production ability, humor styles, and professional success at comedy are three different variables, and the study measured all three because they do not reduce to one another.
10 The Sex Difference and the Meta-Analysis About It
The one meta-analysis in this area is about sex differences in humor production rather than about humor and intelligence, and the distinction is worth stating because the two are often confused. Greengross, Silvia, and Nusbaum published it in Journal of Research in Personality in 2020, volume 84, article 103886. A random effects model across the available studies produced a combined effect size of d = 0.321, with men's humor output rated funnier on average by judges.
An effect of about a third of a standard deviation is a real difference in means and a poor basis for any judgment about an individual. Two distributions separated by d = 0.32 overlap across roughly 87 percent of their range. Picking one man and one woman at random, the man's rated humor would be higher about 59 percent of the time, which is barely distinguishable from a coin flip. Population level effects of this size are genuine and individually invisible, which is the same arithmetic that governs every group comparison in the ability literature.
The meta-analysis has been publicly criticized on methodological and conceptual grounds, including how the underlying studies defined and scored humor production. That criticism does not make the pooled estimate disappear, and it does not settle the question either. What it means is that a reader should treat d = 0.321 as the best available pooled figure from a contested literature rather than as a fixed constant.
A source this page could not verifyA meta-analysis by Greengross and colleagues in 2020 in Psychology of Aesthetics, Creativity, and the Arts on humor production ability and intelligence is sometimes referenced. No such paper could be located. The 2020 Greengross, Silvia, and Nusbaum meta-analysis is on sex differences and appeared in Journal of Research in Personality. If a pooled estimate of the humor and intelligence correlation exists in the peer reviewed record, this page has not found it, and the individual studies in the table above remain the best evidence available.
The primary studies show the same sex pattern in the same direction. Greengross and Miller reported d = 0.38 in their 400 person sample. Howrigan and MacDonald reported that general intelligence did not interact with participant sex in predicting rated humor, meaning the ability association held equally in both halves of their sample even where the mean differed.
11 The Mechanism Under the Joke
Humor comprehension has a documented two stage structure, and both stages draw on capacities that cognitive tests sample directly. Vrtička, Black, and Reiss reviewed the imaging evidence in Nature Reviews Neuroscience in 2013, volume 14, pages 860 to 868. Their account describes a core network in which temporo-occipito-parietal regions detect and resolve incongruity, the mismatch between what was expected and what arrived, after which mesocorticolimbic reward structures and the amygdala produce the response people experience as mirth.
Detection and resolution are the cognitive half, and they map onto measurable operations. Detecting an incongruity requires having built an expectation, which means holding the setup in working memory while the punchline arrives. Resolving it requires finding a second interpretation under which the incongruity makes sense, which means searching semantic memory for an alternative frame. That search is exactly the retrieval operation Christensen and colleagues measured with fluency tasks and found to carry a unique effect.
Semantic retrieval. Finding the second meaning requires searching stored knowledge quickly and selectively, which is what broad retrieval ability measures and what vocabulary tests partly index.
Incongruity detection. Noticing the mismatch requires an expectation to have been formed and held, which is a working memory operation.
Verbal fluency and compression. Getting the collision into a single line under time pressure draws on the same machinery a verbal reasoning subtest samples.
Selection. Most candidate jokes are bad, so producing a good one requires generating several and discarding most, which is a control operation rather than a knowledge one.
This is why the ability association shows up where it does. A test that samples verbal knowledge and working memory is sampling two of the four operations above, which is enough to produce a correlation around .4 and not enough to produce one around .8. The specific subtests that load on working memory are described on the Digit Span page and the Alphanumeric Sequencing page, and how memory relates to measured ability more broadly is on the memory page.
Nothing in the imaging account licenses a causal claim in either direction. It describes what the brain does while processing a joke, not what makes one person better at it than another. Correlation between rated funniness and vocabulary is consistent with shared verbal machinery, with shared exposure to language, and with practice at both, and the studies here cannot separate those.
12 What a Cognitive Test Can and Cannot Tell You Here
No cognitive battery measures humor, and this one does not either. What a battery can do is measure the abilities that the humor literature identifies as carrying the association. That is a narrower and more useful claim than the one the popular coverage makes.
The ability most implicated is verbal knowledge. ACIS reports it as the Verbal Comprehension Index, built from five subtests: Antonyms, Vocabulary, Information, Synonyms, and Similarities. In the published technical manual the index has an omega reliability of .9745, a standard error of measurement of 2.40 points, and a g loading of .864. The subtest g loadings within it run from .865 for Similarities and .860 for Antonyms down to .828 for Synonyms and .814 for Information. Fluid reasoning is reported separately as the FRI, with an omega of .9727 and a g loading of .922, and the two indices correlate .801 with each other, which is why separating them requires modeling rather than eyeballing.
That .801 figure is the reason the Christensen bifactor result matters so much. When two indices correlate that highly, a raw correlation between either one and an outcome is mostly carrying shared general variance. Only a model that separates the general factor from the specific ones can say whether the verbal part is doing independent work, and in the humor case it turned out that the fluid part was not. The full index structure and the six CHC domains are described on the CHC model page and on the cognitive domains page, with the reliability machinery explained on the reliability and validity page.
What a result does not tell you is whether you are funny. A high VCI is consistent with the profile the humor production studies describe and predicts nothing about any individual's ability to write a caption, because a correlation of .49 leaves roughly three quarters of the variance unexplained. ACIS is an unsupervised online assessment, not a clinical instrument, and it is not appropriate for diagnosis, hiring decisions, accommodations, or society admission.
Every number on this page comes from one of the papers below, and where a figure was taken from a later paper's report of an earlier one rather than from the original, the text says so. That distinction matters in a small literature where the same handful of correlations get passed forward across citations.
Greengross, G., and Miller, G. (2011). Humor ability reveals intelligence, predicts mating success, and is higher in males. Intelligence, 39, 188 to 192. N = 400, r = .38 for vocabulary and .25 for Raven's, z = 2.1 for the difference, d = 0.38 for the sex difference.
Christensen, A. P., Silvia, P. J., Nusbaum, E. C., and Beaty, R. E. (2018). Clever People: Intelligence and Humor Production Ability. Psychology of Aesthetics, Creativity, and the Arts, 12, 231 to 241. N = 270, latent correlations of .49, .38, and .22, and a bifactor model in which the fluid effect falls to .09. A copy of the accepted manuscript is available as a PDF.
Howrigan, D. P., and MacDonald, K. B. (2008). Humor as a Mental Fitness Indicator. Evolutionary Psychology, 6, 652 to 666. N = 185, r = .29, regression beta of .20 for intelligence and .16 for extraversion. The full text is available as a PDF.
Kellner, R., and Benedek, M. (2017). The Role of Creative Potential and Intelligence for Humor Production. Psychology of Aesthetics, Creativity, and the Arts, 11, 52 to 58. 151 participants writing punch lines for six cartoons. The correlations of .37 and .13 quoted here are as reported by Christensen and colleagues in 2018.
Willinger, U., and colleagues (2017). Cognitive and emotional demands of black humour processing. Cognitive Processing, 18, 159 to 167, open access at PubMed Central. N = 156, three clusters, group means of 109.7 verbal and 118.1 nonverbal in the highest group.
Silvia, P. J., Greengross, G., Cotter, K. N., Christensen, A. P., and Gredlein, J. M. (2021). If You're Funny and You Know It. Journal of Research in Personality, 92, 104089. Seven studies, n = 1,133, self to judge correlation of .13.
Greengross, G., Silvia, P. J., and Nusbaum, E. C. (2020). Sex differences in humor production ability: A meta-analysis. Journal of Research in Personality, 84, 103886. Combined d = 0.321.
Greengross, G., Martin, R. A., and Miller, G. (2012). Personality traits, intelligence, humor styles, and humor production ability of professional stand-up comedians compared to college students. Psychology of Aesthetics, Creativity, and the Arts, 6, 74 to 82. 31 comedians against 400 students.
Vrtička, P., Black, J. M., and Reiss, A. L. (2013). The neural basis of humour processing. Nature Reviews Neuroscience, 14, 860 to 868. The incongruity detection and resolution account.
Feingold, A., and Mazzella, R. (1991). Psychometric intelligence and verbal humor ability. Personality and Individual Differences, 12, 427 to 435, and Masten, A. S. (1986). Both are quoted here as reported by Howrigan and MacDonald in 2008 rather than from the originals.
Two limits belong on the record. The samples behind almost every figure above are university students in the United States, which restricts the range of both ability and age and inflates nothing but does narrow generalization. And no study cited here is longitudinal or experimental, so nothing on this page establishes that verbal ability causes funniness rather than the two sharing a common source. Correlational evidence supports a description of who tends to be funnier, not a mechanism for making anyone funnier.
The professional framework for interpreting any of this is explicit. The International Test Commission guidelines on test use and the Standards for Educational and Psychological Testing (2014), published jointly by the American Educational Research Association, the American Psychological Association, and the National Council on Measurement in Education, require that score interpretation be tied to documented evidence for the specific use, that classification be reported alongside its measurement error, and that the limits of the reference sample be disclosed. APA standards on test use impose the same discipline on anyone reporting a result: state what the evidence supports, state what it does not, and never let a correlation about a group become a statement about a person.
14 Frequently Asked Questions
Are funny people actually more intelligent?
People who produce funnier material do score higher on average. Christensen and colleagues reported in 2018 that humor production ability correlated .49 with crystallized ability across 270 adults, with a general factor predicting it at .51. The relationship is real and leaves most of the variance unexplained.
How is humor ability measured in these studies?
Participants write captions for cartoons or responses to prompts, and independent judges rate the funniness without knowing who wrote what. Greengross and Miller used six judges across three cartoons, with interrater agreement averaging 0.72.
Does the research use self-reported funniness?
No, and for good reason. Silvia and colleagues found in 2021 that across seven studies of 1,133 adults, people's ratings of their own humor attempts correlated only .13 with the ratings judges gave those same attempts.
Is the humor link with verbal ability or with reasoning?
Mostly verbal. Greengross and Miller reported .38 for vocabulary against .25 for Raven's matrices and confirmed the difference with a Fisher transformation at z = 2.1. Christensen and colleagues found .49 for crystallized ability against .22 for fluid.
What happened to the fluid reasoning effect in the bifactor model?
It vanished. Christensen and colleagues found that once a general factor was modeled, fluid reasoning's unique effect on humor production fell to .09 with p = .398. The variance in the fluid tasks that predicted humor was general variance, not fluid variance.
Does dark humor mean higher intelligence?
The study behind that headline measured comprehension and preference for 12 cartoons in 156 people, not the ability to be funny. The group that liked them most averaged 109.7 on a vocabulary test and 118.1 on a speeded number connection task, which are group means, not individual descriptions.
What was wrong with the coverage of the black humor study?
Three things. It treated an appreciation study as a production study, it reported High Average group means as though they described gifted individuals, and it called a speeded number connection task a measure of nonverbal intelligence when it sits much closer to processing speed.
What separated the groups most in the black humor study?
Mood, not ability. The group with the lowest preference for black humour averaged 19.5 on mood disturbance against 9.2 and 8.6 for the other two, and 14.5 on aggressiveness toward others against 11.2 and 9.4.
Is appreciating humor as strongly linked to ability as producing it?
No. Every effect above .20 in this literature comes from a production study. Appreciation ratings are compressed because most people find most professional material at least somewhat funny, which leaves less variance for ability to explain.
How large is the sex difference in humor production?
Greengross, Silvia, and Nusbaum reported a pooled effect of d = 0.321 in their 2020 meta-analysis, with men's output rated funnier on average. Two distributions separated by that much overlap across about 87 percent of their range.
Does that mean men are funnier than women?
It means the group means differ by about a third of a standard deviation in judged output. Picking one person from each group at random, the man's rating would be higher about 59 percent of the time, which is not far from chance.
Do professional comedians score higher on ability tests?
In the one study that looked, yes. Greengross, Martin, and Miller compared 31 professional stand-up comedians with 400 students in 2012 and found the comedians higher on verbal intelligence, humor production, and all four humor styles.
How much weight should the comedian study carry?
Limited weight on its own. Thirty-one people supports a group comparison but not precise within-group correlations, and comedians are a survivorship sample selected on years of practice as well as on ability.
What is broad retrieval ability and why does it matter here?
It is the ability to search semantic memory quickly, measured with tasks like naming as many animals as you can in a fixed time. It correlated .38 with humor production and kept a unique effect of .20 after general ability was modeled, the only specific factor that did.
Does ACIS measure retrieval fluency?
No. The 20 subtests cover verbal comprehension, fluid reasoning, quantitative reasoning, visual spatial processing, working memory, and processing speed, and none of them is a fluency task. The retrieval effect in the humor literature has no direct counterpart in an ACIS profile.
What happens in the brain when someone gets a joke?
Vrticka, Black, and Reiss described a two stage process in 2013: temporo-occipito-parietal regions detect and resolve the incongruity, then mesocorticolimbic reward structures and the amygdala produce the mirth response.
Are people accurate about which of their own jokes worked?
Slightly. The within-person correlation between self and judge ratings was .13 with an interval from .07 to .19. Notably, people were modest rather than inflated when rating specific attempts, which is the opposite of the usual funnier-than-average pattern.
Does personality predict being funny?
Extraversion does, weakly. Howrigan and MacDonald found a standardized beta of .16 for extraversion against .20 for general intelligence in a regression with ten predictors, while openness, conscientiousness, agreeableness, and neuroticism showed no unique effect.
Can any of this establish causation?
No. Every study cited here is correlational and cross-sectional, none is experimental or longitudinal, and none can separate verbal ability causing funniness from the two sharing a common developmental source.
Which ACIS index is closest to the ability implicated here?
The Verbal Comprehension Index, built from Antonyms, Vocabulary, Information, Synonyms, and Similarities. In the technical manual it has an omega reliability of .9745, a standard error of measurement of 2.40 points, and a g loading of .864.
Will a high score tell me whether I am funny?
No. A correlation of .49 leaves roughly three quarters of the variance in rated funniness unexplained, and ACIS contains no humor task at all. A verbal score describes performance on verbal subtests, nothing more.
Take the assessment
You get a profile, not a number
ACIS measures six CHC domains across 20 subtests and reports each one with its own normed score and confidence interval, so you can see where you are strong and where you are not.