Cannabis and IQ: what the replication record actually shows
One cohort study reported that persistent adolescent onset cannabis users lost eight IQ points by age 38. Within months a simulation reproduced the same pattern with no causal effect at all, and then the twin studies arrived. This page walks the whole sequence in order, with each design and its specific weakness.
A finding is not settled by its first publication. It is settled by what happens to it afterwards.
0 The Short Answer
Heavy cannabis users score lower on intelligence tests than non users, and that gap is real and repeatedly measured. What remains unresolved is whether cannabis caused any of it. Every study design that can separate cause from background has been pointed at this question, and the ones that control family background most tightly find the least. That is the whole story in two sentences, and it is not the story most coverage tells.
The honest complication runs in both directions. The single most cited result, from a New Zealand birth cohort followed since 1972, is genuinely well designed: it measured intelligence years before anybody in the sample had touched cannabis, which is exactly the feature almost no other study has. But the headline figure rests on 23 people, and a critique published the following year showed that an ordinary socioeconomic confounder, with no causal effect of cannabis whatsoever, reproduces the same numbers. The four large twin studies that followed mostly found nothing once siblings were compared with each other. Then the same cohort reported again at age 45 and still found a deficit.
So the useful output of this page is not a verdict. It is a method. If you read to the end you will know how a scientific controversy gets tested, which is more durable than any answer about cannabis, and it transfers to every other claim you will read about what an intelligence test actually measures.
8 points
Average IQ loss from age 13 to age 38 among adolescent onset persistent users in the Dunedin cohort, reported by Meier and colleagues in 2012.
23 people
The size of the subgroup carrying that headline figure, inside a cohort of 1,037.
Near zero
Difference in IQ change between twins discordant for cannabis use in Jackson and colleagues, 2016, across two independent samples.
Not medical adviceThis page describes published research and the methods behind it. It is not medical, psychiatric or legal advice, it is not a substance use assessment, and it cannot tell you anything about your own brain. If cannabis use is affecting your life, that is a conversation for a clinician, not a web page.
1 Three Questions Hiding Inside One
The phrase "does weed lower your IQ" bundles together at least four separate questions with four separate evidence bases, and the arguing usually happens because two people are answering different ones. Untangling them first makes the rest of this page short.
The first is acute impairment. While a person is intoxicated, performance on attention, memory and reaction tasks drops. Nobody disputes this and no serious researcher argues about it. It is the same category of claim as saying that alcohol impairs driving, and it says nothing about lasting capacity.
The second is residual effect. For some period after use stops, performance may remain below baseline, through withdrawal, disturbed sleep or lingering metabolites. This is where a surprising amount of the literature actually sits, because many studies tested people who had used within the previous day or two. Since sleep quality alone moves test scores, as covered in the evidence on sleep and test performance, a study that does not enforce abstinence is measuring something other than durable ability.
The third is persistent change: has the person's capacity been altered in a way that survives weeks or years of abstinence? This is the question people mean when they say cannabis lowers IQ, and it is the hardest one to answer.
The fourth cuts across all three: does the age at which use began matter? Adolescent onset and adult onset are different exposures to a differently developed brain, and the evidence for them is not the same evidence. Treating them as one question is the most common error in popular coverage of this topic.
Question
What would count as evidence
State of the evidence
Acute impairment during intoxication
Laboratory dosing studies with placebo control
Settled. Impairment is well documented and uncontroversial.
Residual effect in the days after use
Testing at controlled intervals after last use
Substantial, and it shrinks sharply once abstinence passes roughly 72 hours.
Persistent change after adolescent onset use
Pre exposure baseline plus a design that controls family background
Contested. This is the entire argument.
Persistent change after adult onset use
Same, in people who began after brain maturation
Weak signal. The cohort that found adolescent effects found none here.
Everything that follows concerns the third row, with the fourth used as a control condition. Keep the rows separate and the literature stops looking contradictory.
2 Step One: The Dunedin Finding of 2012
In 2012, Meier and colleagues published in the Proceedings of the National Academy of Sciences, volume 109, issue 40, pages E2657 to E2664, a result that reframed the whole field: persistent cannabis users showed neuropsychological decline from childhood to midlife, and the decline was concentrated in those who started as adolescents. The full text is open access in PubMed Central.
The data came from the Dunedin Multidisciplinary Health and Development Study, a birth cohort of 1,037 children born in Dunedin, New Zealand, between April 1972 and March 1973, representing 91 percent of eligible births and constituted at age 3. Members were assessed at ages 5, 7, 9, 11, 13, 15, 18, 21, 26, 32 and 38. At the age 38 wave, conducted between 2010 and 2012, 96 percent of the 1,004 living members took part. Cannabis use and dependence were assessed by interview at 18, 21, 26, 32 and 38. Intelligence was measured in childhood with the Wechsler Intelligence Scale for Children Revised and again at 38 with the WAIS-IV, the predecessor of the edition described in what changed between the fourth and fifth Wechsler editions, both on the familiar scale where the mean is 100 and the standard deviation is 15.
The core result was a dose response. Compared with their own childhood scores, members with one, two, or three or more cannabis dependence diagnoses across those waves showed IQ changes of minus 0.11, minus 0.17 and minus 0.38 standard deviation units. The paper spells out what the largest of those means in points: a fall from 99.68 to 93.93, roughly six IQ points. Members who never used cannabis showed a slight increase.
Then came the finding that travelled. Splitting by age of onset, the decline sat almost entirely with those who began in adolescence. Adolescent onset users who kept using persistently through to age 38, a group of 23 people, lost an average of eight IQ points from age 13 to age 38. Adult onset persistent users showed no comparable decline. Cessation did not fully restore functioning in the adolescent onset group. Third party informants, people nominated by the study members as knowing them well, independently reported more attention and memory problems in persistent users, which is a genuinely useful corroboration because it does not depend on the test at all.
The authors were explicit that the pattern was suggestive of a neurotoxic effect on the developing brain. That word, suggestive, did most of the work they intended it to do and almost none of the work the coverage gave it.
3 Why That Cohort Was Worth Taking Seriously
Most research linking a substance to lower cognitive scores cannot tell you whether the scores were already lower before exposure began. Dunedin can, and that single property is why the 2012 paper deserved the attention it got. Understanding why makes the later critiques legible instead of looking like point scoring.
Consider the ordinary cross sectional study: recruit heavy users and non users, test both, compare. If users score lower, you have learned nothing about causation, because the people who become heavy users differ from those who do not in dozens of measured and unmeasured ways before they start. The comparison is contaminated at the moment of recruitment.
Dunedin sidesteps that by measuring the outcome first. Intelligence was assessed at ages 7, 9, 11 and 13, when only seven members of the entire cohort reported having tried cannabis. The analysis then used change from that baseline rather than the adult level, so a group that started lower does not automatically look damaged. That structure is called a difference in differences estimator, and it removes any confounder whose effect on ability was already complete before age 13 and stayed constant afterwards.
Three further design properties matter. The cohort is population representative rather than clinic recruited, so it is not selected on pathology. It is large and long enough that the cohort effectively forms its own reference frame, which sidesteps most of the trouble described in how norms are built. And attrition was not differential: the four percent who missed the age 38 wave were no more likely to have been cannabis dependent at 18 than those who attended, which rules out the most obvious selective dropout story.
What the design still could not doNobody was randomly assigned to use cannabis, so unmeasured differences between groups remain possible. Exposure was self reported in retrospective interviews, not biologically verified, and it was counted in dependence diagnoses rather than in quantity of THC. And the headline subgroup contained 23 people, which means one unusual individual moves the average by a noticeable amount.
Those three limits are not nitpicks invented by critics after the fact. They are the standard weaknesses of observational cohort work, and the authors named most of them. What happened next is that somebody took the second and third seriously enough to build a model.
4 Step Two: The Confounding Model That Reproduced the Pattern
In January 2013, the economist Ole Rogeberg published in the same journal, PNAS volume 110, issue 11, pages 4251 to 4254, a paper showing that a plausible socioeconomic confounder, containing no causal effect of cannabis at all, generates the Dunedin pattern. This is the step that popular coverage almost universally omits, and it is the most instructive step in the sequence. The paper is available in full through PubMed Central.
Rogeberg's argument starts from the exact property that made the Dunedin design strong. A difference in differences estimator is unbiased only if omitted variables have no time varying effect on the outcome. So the question becomes: is there a variable that predicts adolescent cannabis use and also changes a person's IQ trajectory between 13 and 38? He nominated socioeconomic status, and gave a mechanism drawn from the Flynn and Dickens model of reciprocal causation between ability and environment, the same body of theory that sits behind the secular rise in measured scores.
The mechanism runs like this. Compulsory schooling is a cognitively demanding environment imposed on everyone regardless of background, and it lifts the measured scores of low socioeconomic status children more than those of high status children, because the increment over their home environment is larger. When school ends, people increasingly select their own environments, and the non cognitive factors that correlate with background push some of them toward less cognitively demanding settings. The temporary boost fades. If cannabis use is also more common in that group, you get an association between use and apparent decline without any pharmacology involved.
He then simulated it. The model assumed only that socioeconomic status predicts cannabis exposure, and that lower status groups receive a temporary schooling boost averaging about four IQ points, roughly a quarter of a standard deviation, which unwinds with age. Across 500 simulated runs of 875 people each, the pattern of effects reproduced the actual Dunedin estimates, which came from about 894 people. His conclusion was precise: the causal effects are likely to be overestimates and the true effect could be zero, and he added that it would be too strong to say the results had been discredited.
He also pointed at a fingerprint inside the original paper. Among members with high school education or less, the most exposed group lost 0.45 standard deviation units, based on 26 people. Among those with more education, the comparable figure was about half that, 0.24 units, based on 12. That is what a background confound looks like, although as Rogeberg conceded, reduced schooling could equally sit on the causal path itself.
5 Step Three: The Reply, and the Reply to the Reply
The Dunedin team answered in the same issue, and their reply is a model of how to respond to a methodological critique: they went back to the data and tested the proposed mechanism directly rather than defending the interpretation. Moffitt, Meier, Caspi and Poulton published at PNAS volume 110, issue 11, pages E980 to E982, readable in PubMed Central.
Rogeberg's model needed two things to be true in the Dunedin cohort. First, that adolescent cannabis users are concentrated in lower socioeconomic groups. Second, that the measured scores of lower status children rise during schooling and fall afterwards. The reply tested both. Low socioeconomic status did not significantly predict adolescent onset cannabis dependence, with a chi square of 1.15 and a p value of 0.56, and only 23 percent of adolescent cannabis users came from low status families. The trajectory claim failed too: scores of children from low status families did not rise from the start of schooling to adolescence and did not fall from adolescence to adulthood, and in the direct test, low status was unrelated to adolescent to adult IQ decline, with a correlation of minus 0.006 and a p value of 0.86.
They then did the two things that settle an argument about confounding within a fixed dataset. Controlling statistically for socioeconomic status left the association essentially unchanged, from a standardised coefficient of minus 0.152 to minus 0.158. Restricting the analysis to members from middle class families only, excluding both the low and high ends so that low status confounding could not operate, also left it standing, at minus 0.155 with a t of minus 3.61 and a p value of 0.0003.
A second commentator, Daly, proposed in the same issue at page E979 that conscientiousness, one of the five factor traits discussed in personality traits and measured ability, might explain the association. The Dunedin team had a pre exposure measure of childhood self control, a precursor of that trait, and it was unrelated to IQ decline, with a correlation of minus 0.034. Controlling for it changed nothing. Rogeberg published a short rejoinder at page E983, titled to the effect that causal inference from observational data remains difficult, and there the exchange ended.
The lesson is not that one side won. Both sides argued well, and neither could settle it, because they were arguing about a single dataset whose confounding structure is fixed. That is precisely the situation in which a different design, not a better argument, is the only way forward. Which is what happened next.
6 Step Four: The Twin Design
In 2016, Jackson and colleagues published in PNAS, volume 113, issue 5, pages E500 to E508, the study that changed most researchers' priors: two independent longitudinal twin samples, 789 and 2,277 adolescents, tested before and after cannabis involvement. The full paper is in PubMed Central.
The logic of a discordant twin design is worth stating plainly, because it is the single most useful research tool in this entire debate. Within a twin pair, both members share a household, a neighbourhood, a school district, parental education and parental income. Identical twins additionally share their entire genome. If one twin uses cannabis and the other does not, and the user declines more than the co twin, the family background explanations are eliminated in one move rather than modelled one variable at a time. If the user does not decline more, the association measured across unrelated people was probably driven by the things twins share.
The two samples were deliberately different: the Risk Factors for Antisocial Behavior study in Southern California, predominantly Hispanic, tested at ages 9 to 10 and followed to 19 to 20, and the Minnesota Twin Family Study, largely white, tested at 11 to 12 and followed for seven years. Cognitive assessment at both points used Wechsler subtests including Vocabulary, Information, Similarities, Block Design, Matrix Reasoning and Picture Arrangement.
Across unrelated participants, the association replicated. Users scored lower, and declined relative to non users on the verbal subtests, mainly the Vocabulary subtest and general information, at a magnitude of just under four IQ points in both samples independently. That is a genuine finding and close to what the Dunedin numbers imply, so the twin studies did not fail to see what everyone else saw.
Then came the two tests that mattered. There was no dose response: participants who used daily for six months or longer showed no greater change than those who had tried cannabis fewer than 30 times, in either sample, before or after adjusting for other substance use. And in the co twin comparison there were no significant differences in IQ change on Vocabulary, Similarities, Information, Block Design or Picture Arrangement, in identical or fraternal pairs. The one statistically significant result went the wrong way: on Matrix Reasoning, the using twin improved more than the abstinent one. Restricting to the 47 Minnesota pairs where one twin had never used and the other was a heavy user, the strongest available test, produced nothing.
The power caveat, stated honestlySome of these subgroups are small. Only 25 identical discordant pairs were available in the California sample. A study that finds no difference in a small sample has not shown the difference is zero, it has failed to detect it, and those are different claims. The force of this result comes from two independent samples agreeing, and from the missing dose response, not from any single null.
7 Step Five: Three More Genetically Informative Samples
A single twin study can be a fluke. Three more arrived within five years, in different countries with different cohorts, and they converged. The most striking one was led by the first author of the original Dunedin paper.
Meier and colleagues published in Addiction, volume 113, issue 2, pages 257 to 265, in 2018, a co twin control analysis of 1,989 twins from the Environmental Risk study, a nationally representative birth cohort born in England and Wales in 1994 and 1995, with IQ measured at ages 5, 12 and 18. The findings are worth reading twice. Adolescents who met criteria for cannabis dependence scored 5.61 points lower at age 12, before initiation, and 7.34 points lower at 18. But the change from 12 to 18 was not significantly different, with a t of minus 1.27 and a p value of 0.20. Executive function differences between users and non users largely disappeared within twin pairs, on five of six tests. The paper is open in PubMed Central, and its conclusion states that family background factors explain why adolescent cannabis users perform worse.
Ross and colleagues published in Drug and Alcohol Dependence, volume 206, article 107712, in 2020, a quasi experimental co twin analysis from the Colorado longitudinal twin samples covering adolescence into the twenties. Phenotypic associations between cannabis and cognition existed but became negligible after accounting for other substance use, and within twin pairs almost nothing survived, with one exception: cannabis frequency at 17 predicted worse general cognitive ability and executive function at 23. Their own summary, available in PubMed Central, is that they found little support for a causal effect.
Schaefer and colleagues published in PNAS, volume 118, issue 14, article e2013180118, in 2021, pooling three longitudinal twin studies for 3,762 participants. In identical co twin comparisons, which absorb genetic and shared environmental confounding completely, the cognitive and psychiatric associations dropped out. What survived was socioeconomic: educational attainment, occupational status and income. Follow up analysis pointed at a mechanism, since greater use during the school years tracked with falling grade point average, lower academic motivation and more disciplinary problems. That result, readable in PubMed Central, is the most nuanced in the literature: a plausible causal effect on life trajectory alongside little evidence of one on ability. It sits alongside what is known about school achievement correlations and the association between schooling and measured ability without contradicting either.
Four genetically informed designs, four failures to find a within pair cognitive effect of the size the cross sectional literature implies. That is a pattern, not a coincidence.
8 Step Six: What Pooling the Studies Actually Adds
Meta-analysis is often treated as the top of the evidence pyramid, but pooling a hundred confounded studies gives you a very precise estimate of a confounded quantity. The value of the meta-analytic work here is not the headline number, it is the moderator analysis, which quietly reveals what much of the field had been measuring.
Scott and colleagues published in JAMA Psychiatry, volume 75, issue 6, pages 585 to 595, in 2018, the first quantitative synthesis for adolescents and young adults: 69 studies, 2,152 cannabis users with a mean age of 20.6 and 6,575 comparison participants. The overall effect for frequent or heavy use was small, a standardised mean difference of minus 0.25 with a confidence interval from minus 0.32 to minus 0.17. Then the crucial split. Fifteen studies that required abstinence longer than 72 hours before testing, covering 928 people, produced an effect of minus 0.08 with a confidence interval from minus 0.22 to plus 0.07, which is not distinguishable from zero. The 54 studies with looser abstinence criteria produced minus 0.30. The authors concluded that previous work may have overstated the magnitude and persistence of deficits, and that reported deficits may reflect residual effects of recent use or withdrawal. The paper is open in PubMed Central.
That moderator result is the second most important finding on this page, after the twin studies. It says that a large share of the cannabis cognition literature was measuring how recently someone had used, not what they were capable of. Anyone who has read why subtests are timed or how to prepare on the day already knows that state factors move scores by several points, which is exactly the magnitude in play here. A whole literature can rest on that substitution rather than a share of one: the three studies carrying the claim that generative artificial intelligence lowers intelligence recorded EEG connectivity, survey answers and perceived mental effort, so the AI case contains no ability measure at all.
A second confounder keeps appearing. Mokrysz and colleagues, in the Journal of Psychopharmacology, volume 30, issue 2, pages 159 to 168, in 2016, analysed 2,235 teenagers from the Avon Longitudinal Study of Parents and Children. After full adjustment, those who had used cannabis 50 times or more did not differ from never users on IQ at 15 or on exam performance at 16. Adjusting for cigarette smoking dramatically attenuated the associations, and tobacco use predicted educational outcomes robustly even with cannabis users excluded. That paper is also in PubMed Central.
The most recent synthesis narrows the target. Pilon, Dumais, Giguere and Potvin, in Addictive Behaviors, volume 170, article 108434, in 2025, restricted the question to people diagnosed with cannabis use disorder rather than users in general. Across 13 cognitive domains they found small to moderate impairments in 10, with verbal learning and memory, processing speed and working memory largest at around 0.4 to 0.5 standard deviations. Those are cross sectional clinical comparisons, so direction of causation remains open, but the finding is a fair warning against writing the whole question off.
9 Step Seven: The Same Cohort, Seven Years Later
In 2022 the Dunedin team reported again, this time at age 45, and the deficit was still there. Meier and colleagues published in the American Journal of Psychiatry, volume 179, issue 5, pages 362 to 374, with 94 percent retention in a cohort now followed for four and a half decades. The paper is open in PubMed Central.
Cannabis use and dependence had been assessed at 18, 21, 26, 32, 38 and 45. Intelligence was measured at 7, 9, 11 and again at 45. Long term users showed a mean decline of 5.5 IQ points from childhood to midlife, poorer learning and processing speed relative to their own childhood scores, and informant reported memory and attention problems. Long term users also had smaller hippocampal volume, though hippocampal volume did not statistically mediate the cognitive deficits, which is an important negative: the obvious anatomical explanation did not carry the effect.
The strongest feature of this paper is not the deficit, it is the specificity analysis. The same comparisons were run for long term tobacco users, long term alcohol users, midlife recreational cannabis users and cannabis quitters. In those groups the deficits were absent or smaller. Using alcohol as the comparison exposure is a sharper move than it looks, because that literature carries a pre exposure baseline of its own: in a Danish conscript follow up, the men who later acquired an alcohol related hospital diagnosis were already 5.5 points below everyone else at conscription, so the premorbid gap recorded before any diagnosis existed is documented in the control substance rather than assumed away. The cognitive shortfall also survived adjustment for persistent tobacco, alcohol and other illicit drug use, for childhood socioeconomic status, for low childhood self control and for family history of substance dependence. Each of those adjustments individually answers a specific alternative explanation, and collectively they are the most thorough confound accounting any observational study in this area has done.
It does not close the case, for two reasons that should be said out loud. First, it is the same cohort and largely the same analysts, so it is a continuation, not an independent replication. Second, it is still observational, and a design that cannot rule out unmeasured confounding at 38 cannot rule it out at 45 either. What it does establish is that whatever is going on in this cohort is durable, dose linked, and not shared with tobacco or alcohol.
Laid out in one table, the sequence stops being a fight and becomes a fairly ordinary case of a strong association surviving statistical adjustment but not surviving genetic control. Each row names a design and what it could and could not establish.
Study
Design
Sample
Result
Meier and colleagues, 2012, PNAS
Birth cohort, IQ at 13 and 38
1,037 born 1972 to 1973
8 point loss in adolescent onset persistent users, n equals 23; dose response by dependence diagnoses
Rogeberg, 2013, PNAS
Simulated socioeconomic confounding model
500 runs of 875
Reproduced the same pattern with no causal effect; true effect could be zero
Moffitt and colleagues, 2013, PNAS
Reanalysis of the same cohort
Same cohort
Low status unrelated to decline, r equals minus 0.006; association held in middle class only subsample
Jackson and colleagues, 2016, PNAS
Two twin cohorts, co twin control
789 and 2,277
Four point verbal gap across unrelated people; nothing within discordant pairs; no dose response
Meier and colleagues, 2018, Addiction
Co twin control, E-Risk cohort
1,989 twins
Users already lower at age 12; no greater decline from 12 to 18
Scott and colleagues, 2018, JAMA Psychiatry
Meta-analysis of cross sectional studies
69 studies, 8,727 people
Overall d minus 0.25; d minus 0.08 when abstinence exceeded 72 hours
Ross and colleagues, 2020, Drug Alcohol Depend
Co twin control, Colorado samples
Longitudinal twin pairs
Little support for a causal effect; one surviving within pair association
Schaefer and colleagues, 2021, PNAS
Three twin cohorts, identical co twin control
3,762
Cognitive associations dropped out; education, occupation and income effects survived
Meier and colleagues, 2022, Am J Psychiatry
Same birth cohort at age 45
1,037, 94 percent retention
5.5 point decline in long term users, specific to cannabis; smaller hippocampal volume, non mediating
What is agreed: intoxication impairs performance; heavy users score lower than non users; a good part of that gap is present before use begins; residual effects account for much of what remains and fade over days; and cannabis use disorder, as a clinical category, is associated with measurable deficits concentrated in the domains covered by memory and measured ability.
What is open: whether persistent heavy adolescent onset use causes durable change, and if so how large; whether the Dunedin cohort is detecting something real that the twin studies are underpowered to see, or whether Dunedin holds a confounder nobody has named; and whether any of this generalises to current products. Freeman and colleagues, in Addiction, volume 116, issue 5, pages 1000 to 1010, in 2021, pooled 66,747 herbal cannabis samples across eight studies and found THC concentration rising by about 0.29 percentage points per year from 1970 to 2017, with cannabidiol flat. Every cohort described on this page was exposed to a weaker product than is sold today, which means the evidence base is about a substance that has changed underneath it.
11 The Design That Would Settle It
The question is answerable in principle, and naming the study that would answer it is more useful than another round of interpretation. Five features would be needed, and no existing study has all five at once.
Feature
Why it is needed
Who has it
Ability measured before any exposure
Otherwise pre existing differences are read as damage
Dunedin, E-Risk, both twin cohorts in Jackson
Identical twins discordant for use
Removes genes and shared upbringing in one step
Jackson, Meier 2018, Ross, Schaefer
Biologically verified exposure and dose in THC
Self reported occasions do not measure the drug delivered
Essentially nobody, at scale
Enforced abstinence before testing
Otherwise residual effects masquerade as capacity
15 of the 69 studies Scott pooled
Power to detect two to three IQ points
The contested effect is small; nulls in tiny subgroups prove nothing
Only the largest cohorts, and not within discordant pairs
Randomising adolescents to years of cannabis use is unethical and will not happen, so this has to be assembled from natural variation. The realistic routes are a large identical twin registry with prospective childhood testing, verified exposure biomarkers and a supervised abstinence window before retesting; Mendelian randomisation, which uses genetic variants associated with cannabis initiation as instruments and sidesteps confounding if the instrument assumptions hold, a use of genotype quite different from the one explained in what heritability actually means; and staggered legalisation as a natural experiment, which changes exposure at population scale for reasons unrelated to any individual's ability.
Each has a weakness worth naming. Twin registries with all five features are expensive and slow. Genetic instruments for substance use are entangled with personality and risk tolerance, which violates the exclusion restriction that the method needs. Legalisation studies measure population averages, where a real effect concentrated in a small heavy using minority would be diluted to invisibility.
What a null result meansFour twin studies failing to find a within pair effect is meaningful evidence, but it is not proof that the effect is zero. It bounds the effect: if cannabis causes a durable loss of ability in typical adolescent users, that loss is probably smaller than the four point gap seen across unrelated people, and possibly much smaller. Bounding an effect and refuting it are different achievements.
The National Academies of Sciences, Engineering, and Medicine reviewed this territory in their 2017 report on the health effects of cannabis and cannabinoids, available through the NCBI Bookshelf, and their treatment of cognition is a fair model of how to write conclusions with graded confidence rather than a verdict.
12 How to Read a Scientific Controversy
The transferable product of this page is six moves, each of which appeared in the sequence above and each of which works on any contested empirical claim. They are worth more than the answer about cannabis, because you will meet a hundred more arguments shaped like this one.
Find the design before the finding. The first question about any result is not what it found but what comparison produced it. Cross sectional, longitudinal with a pre exposure baseline, discordant twin and randomised are four different levels of claim, and the same effect size means something different at each. This is the same discipline that separates a real from a spurious claim in the material on the common myths about testing.
Ask what the comparison group is. Heavy users versus non users is not a comparison between two versions of the same person. Twins versus their siblings almost is. The closer the comparison group is to a counterfactual version of the exposed person, the more the difference can be trusted.
Try to reproduce the pattern with no causal effect. This is Rogeberg's move and it is the most underused tool in public argument. Before accepting that A caused B, build the most plausible model where it did not and see whether that model also produces the observed numbers. If it does, the data have not yet distinguished the two.
Look for replication with a different confound structure. A bigger version of the same design inherits the same bias, only more precisely. What moves knowledge is a design whose weaknesses are different, which is why the twin studies mattered and why a fourth cohort study would not have.
Find the subgroup carrying the headline. The Dunedin number that entered popular culture came from 23 people. That does not make it wrong. It makes it fragile, and fragility is a fact about a result that belongs in any honest summary. The same habit is what keeps score interpretation sane, as in how rare a given score is at the extremes where samples thin out.
Separate not detected from shown to be absent. Four null twin results constrain the size of a possible effect. They do not prove there is none. Popular coverage collapses these constantly, in both directions, and the collapse is where most bad science writing lives. It is the same error behind the ten percent brain myth and behind confident claims about whether scores can be raised.
Run those six on the next study you see summarised in a headline. Most will fail at the second.
13 What a Test Can and Cannot Tell You About This
No intelligence test, ACIS included, can tell an individual whether cannabis changed their cognition, and it is worth being explicit about why rather than leaving the impression that it might. The reason is structural, not a limitation of any particular instrument.
Detecting a change in a person requires a baseline from before the exposure. Almost nobody has one. Without it, a single score tells you where you stand relative to a reference group, not where you would have stood otherwise. Even with two administrations, the comparison is hard: the ACIS Full Scale composite has a standard error of measurement of about 1.60 IQ points, so a difference of four points between two sittings sits close to what ordinary measurement variation plus practice effects can produce on their own. Group findings of the size argued about on this page are simply not resolvable one person at a time. The interval between two sittings is not a neutral gap either, since raw performance climbs and then falls across a life while rank order holds, correlating .66 from age 11 to age 79 in the Lothian cohorts, so what change looks like with no exposure at all is the baseline a personal decline would have to beat. That is the honest content of the Full Scale composite and of the relationship between a score and its percentile.
What a battery can do is show shape. ACIS reports 20 subtests across six primary domains, with a composite omega of .9886 and a g loading of .958 on a technical analysis set of 2,750 complete records, drawn from an adult reference frame of 3,243 English speaking records aged 16 to 90. If someone's concern is memory and speed specifically, the domains that the cannabis use disorder literature implicates most, then a profile that separates Gwm and Gs from Gc and Gv is more informative than any single number, which is the argument for a report with subtest detail over a bare score. Subtests such as Digit Span, Coding and Symbol Search load on precisely those domains, and fluid and crystallized ability behave differently again.
The limits have to be stated with the same clarity as the figures. ACIS is unsupervised, taken at home without a proctor, which means testing conditions are not controlled. Its reference frame is self selected rather than a census sample, so percentile placement carries more uncertainty than the reliability figures alone suggest. It is not a clinical or diagnostic instrument, and it is not appropriate for diagnosis, hiring, accommodations or high IQ society admission. Anyone whose real question is whether their cognition has changed needs a supervised evaluation with a qualified clinician, ideally with records of prior testing, and the routing logic in the comparison between supervised and unsupervised testing lays out when that threshold is crossed. Researchers who want verified administrations for a study of substance use and cognition can use the research workspace or administer the test to participants through private links.
All of this follows from the same source that governs the rest of the site. The Standards for Educational and Psychological Testing (2014), published jointly by the American Educational Research Association, the American Psychological Association and the National Council on Measurement in Education, require that score interpretations be supported by evidence for the specific use proposed, that reliability and measurement error be reported alongside every score, and that limits of interpretation be stated plainly. APA testing standards and the International Test Commission guidelines on test use say the same thing in different words. Under those standards, an unsupervised online score is evidence about a person's current standing on a set of tasks. It is not evidence about what any substance did to them, and no page on this site will claim otherwise.
Heavy users score lower on average, but the studies best able to separate cause from background mostly fail to find that cannabis produced the gap. Much of the difference is present before use begins, and a further part reflects recent use rather than durable capacity.
What did the Dunedin study actually find?
Meier and colleagues, in PNAS in 2012, reported that members of a New Zealand birth cohort with persistent cannabis dependence declined relative to their own childhood scores, with the largest loss, about eight points, among 23 adolescent onset persistent users followed to age 38.
Why is the Dunedin cohort considered strong evidence?
Because intelligence was measured at ages 7 through 13, before nearly anybody in the cohort had used cannabis. That pre exposure baseline lets researchers analyse change rather than level, which removes the most obvious selection problem in substance use research.
What was wrong with the 2012 finding?
Nothing was shown to be wrong with it. Rogeberg demonstrated in 2013 that a socioeconomic confounding model with no causal effect reproduces the same numbers, which means the data did not distinguish the causal reading from a plausible non causal one.
Did the Dunedin team respond to that critique?
Yes, in the same journal issue. They tested the proposed mechanism directly and found that low socioeconomic status did not predict adolescent cannabis dependence in their cohort and was unrelated to score decline, and the association survived restriction to middle class families.
What is a discordant twin study?
It compares twins from the same pair where one used a substance and the other did not. Identical twins share their genome and their upbringing, so any remaining difference cannot be explained by genetics or family background.
What did the twin studies find?
Jackson and colleagues in 2016, Meier and colleagues in 2018, Ross and colleagues in 2020 and Schaefer and colleagues in 2021 all found the raw association between cannabis and lower scores, and all found it largely disappeared when twins were compared with their own siblings.
Does that prove cannabis has no effect on intelligence?
No. Four null results constrain how large a durable effect could be, they do not demonstrate it is zero. Some of those within pair analyses used small numbers of discordant pairs, which limits how confidently absence can be claimed.
Why do adolescent and adult use get treated differently?
Because the brain is still developing through the teens and early twenties, so the same exposure meets a different biological state. In the Dunedin cohort, adult onset persistent users showed no comparable decline, which is why the two must be reported separately.
How much of the deficit is just recent use?
A large share. Scott and colleagues found in 2018 that studies requiring more than 72 hours of abstinence produced an effect indistinguishable from zero, while studies with looser criteria produced a clearly negative one. Timing of the last use drove much of the literature.
Does cannabis kill brain cells?
That framing is not supported and is not how researchers describe the mechanism. The 2022 Dunedin report did find smaller hippocampal volume in long term users, but volume did not statistically account for the cognitive differences observed in the same people.
Is cannabis use disorder different from ordinary use?
The evidence suggests yes. Pilon and colleagues, in Addictive Behaviors in 2025, restricted their meta-analysis to people with a diagnosis and found small to moderate deficits across most cognitive domains, larger than what general user samples show.
Does today's stronger cannabis change the answer?
It might, and nobody knows. Freeman and colleagues documented in 2021 that THC concentration in herbal cannabis rose steadily from 1970 to 2017. Every cohort described here was exposed to a weaker product than is sold now, so the evidence lags the substance.
Which cognitive abilities are most implicated?
Verbal learning and memory, processing speed and working memory show the largest effects in clinical samples. In the twin cohorts, the raw differences appeared mainly on vocabulary and general knowledge subtests, which are crystallized rather than fluid measures.
Could reduced schooling explain the association?
Partly, and this is where the evidence is most interesting. Schaefer and colleagues found that effects on education, occupation and income survived identical twin controls even though cognitive effects did not, with falling grades and school discipline as the likely route.
What about tobacco as a confounder?
It matters more than most coverage acknowledges. Mokrysz and colleagues found in 2016 that adjusting for cigarette smoking dramatically attenuated the cannabis associations, and that tobacco predicted educational outcomes robustly even when cannabis users were excluded from the analysis.
Can an IQ test tell me if cannabis affected me?
No. Detecting personal change requires a score from before exposure, which almost nobody has, and the differences argued about in this literature are smaller than the ordinary variation between two administrations of the same test.
How precise is a single ACIS score?
The Full Scale composite has a standard error of measurement of about 1.60 IQ points, with composite omega of .9886. That is precise for an unsupervised instrument, and still not precise enough to resolve a four point personal change against practice effects.
Is ACIS appropriate for this question?
It can describe your current profile across six domains, which is more informative than a single number if your concern is memory or speed specifically. It is unsupervised and not diagnostic, so it cannot substitute for a supervised clinical evaluation.
Where should I read the original papers?
Almost all of them are open access in PubMed Central, including the 2012 Dunedin paper, the Rogeberg critique, the reply, the 2016 twin study, the 2018 meta-analysis and the 2022 midlife follow up. Every one is linked in the text above.
What is the single most useful thing to take from this page?
That a controversy is resolved by designs with different weaknesses, not by louder arguments about one dataset. Ask what comparison produced a finding, then ask what could produce the same pattern with no causal effect at all.
Take the assessment
You get a profile, not a number
ACIS measures six CHC domains across 20 subtests and reports each one with its own normed score and confidence interval, so you can see where you are strong and where you are not.