AI and IQ: what the famous studies actually measured
Every viral claim that artificial intelligence is lowering human intelligence rests on three studies. None of them measured IQ. One is still an unreviewed preprint of 54 people, and two rely on what participants said about themselves. Here is what each actually tested, and what the evidence will and will not support.
Three different outcomes, three different studies, and none of them an intelligence test.
0 The Short Answer
No published study has shown that generative artificial intelligence lowers IQ, because no published study has measured IQ before and after AI use. That is not a quibble about wording. It is the single most important fact about this entire literature, and almost no article that repeats the claim mentions it.
Three papers carry nearly all of the coverage. Kosmyna and colleagues at the MIT Media Lab recorded electroencephalography from 54 people writing essays, and measured neural connectivity plus the ability to quote one's own essay afterwards. Gerlich, publishing in Societies in 2025, surveyed 666 people about how often they use AI tools and scored them on a critical thinking measure at a single point in time. Lee and colleagues at Microsoft Research and Carnegie Mellon, publishing at the CHI 2025 conference, asked 319 knowledge workers how much mental effort they felt they were expending. Connectivity, survey responses, perceived effort. Not one of those is a measure of general cognitive ability.
The honest complication is that this does not mean nothing is happening. Cognitive offloading is real, it is well documented, and it predates computers by several thousand years. The best study in the field, a randomized controlled trial published in PNAS in 2025, does find a genuine cost, though it measures high school mathematics rather than what an intelligence test measures, and it finds that the cost disappears when the tool is designed differently. The gap between that finding and the headline is where this article lives.
Zero
Published studies measuring IQ before and after generative AI use.
54
Participants in the MIT Media Lab EEG study that generated most of the coverage.
1975
Birth cohort at which measured Norwegian ability scores began falling, decades before generative AI.
Trace almost any article about AI and intelligence back far enough and it terminates in the same three citations. They are cited together, as though they converge on a finding, when in fact they use three incompatible methods on three different populations to measure three different things. Laying them side by side is the fastest way to see the problem.
Study
Sample
Design
What it measured
Peer reviewed
What it cannot tell you
Kosmyna et al. 2025, MIT Media Lab, Your Brain on ChatGPT
54 adults, 18 in the final session
Between subjects across three sessions, one crossover session
EEG connectivity during essay writing, ability to quote one's own essay, sense of ownership
No. Preprint on arXiv, last revised December 2025
Whether any stable ability changed, in either direction
Gerlich 2025, Societies
666 respondents
Cross sectional survey with interviews
Self reported frequency of AI tool use against a critical thinking score
Yes, with a correction published in September 2025
Which came first, the AI use or the score
Lee et al. 2025, CHI conference
319 knowledge workers, 936 task examples
Survey of first hand work examples
Perceived cognitive effort and self reported confidence
Yes, in the CHI 2025 proceedings
Whether actual performance changed at all
Read that table again and notice what is missing from the fourth column. There is no ability test in it. There is no measure of the general factor, no fluid reasoning task, no working memory span, nothing that a psychometrician would recognize as an index of capacity. The studies are not bad for lacking those things. They were not designed to have them. The failure belongs to the coverage, which converted three narrow findings into one broad claim.
What each study is good forKosmyna and colleagues collected a genuinely valuable dataset and built an ambitious multimethod pipeline. Gerlich documented a real association in a large sample. Lee and colleagues described how professionals experience their own thinking while using these tools. Each is a legitimate contribution. None of them is an answer to the question the headlines asked.
2 What the MIT Media Lab Study Actually Did
The study that launched the panic is a 216 page preprint about essay writing, and its central measure is a statistical comparison between electrode pairs. The design is worth stating precisely, because the precision is what the coverage lost. Participants were assigned to one of three conditions: write using a large language model, write using a search engine, or write with no tools at all, described in the paper as the Brain-only group. Each person stayed in the same condition for three sessions. Fifty four people took part, eighteen per group.
During writing, the researchers recorded electroencephalography and derived directed connectivity between electrode pairs. Essays were analyzed with natural language processing and scored by human teachers and by an AI judge. After each session, participants were interviewed, and one of the questions asked them to quote a line from the essay they had just written. Kosmyna and colleagues report that Brain-only participants showed the strongest and most distributed networks, search engine users showed moderate engagement, and language model users showed the weakest connectivity. Language model users also reported the lowest sense of ownership over their essays and struggled to quote their own work. The full preprint is available on arXiv.
Everything in that paragraph is a finding about what happens while you write an essay with a chatbot, and about what you remember of a document you did not compose. That is interesting. It is also close to tautological. If you did not write the sentence, you are less likely to recall the sentence, and you are less likely to feel that it is yours. The interesting question is whether anything persists, and for that you need a different design.
There is a further technical point that is almost universally misread. In this paper a significant connection is defined as a directed coupling between two electrodes that is statistically stronger in one group than in another. It is a comparison, not an absolute quantity. A group with fewer significant connections has not been shown to have weaker overall neural activity, only fewer pairs that beat another group in a statistical test. The phrase weakest connectivity, repeated everywhere, invites a reading the measure does not support. Treating a brain signal as a direct readout of mental power is the same move that keeps the ten percent of the brain myth in circulation, and it deserves the same resistance here. The distance between a brain measure and an ability measure has been quantified: pooled across 88 studies and more than 8,000 people, in vivo brain volume correlates with IQ at .24, about six percent of the variance, which is why the alcohol imaging results count as evidence about brains rather than about scores.
3 The Fourth Session Almost Nobody Reports
The study included a crossover session, and its result is the most informative thing in the paper and the least reported. In session four, the language model group was asked to write with no tools, becoming the LLM-to-Brain group. The Brain-only group was given the language model, becoming the Brain-to-LLM group. Eighteen people completed it, roughly nine per condition.
The paper's own abstract states the outcome plainly. LLM-to-Brain participants showed reduced alpha and beta connectivity, which the authors read as under-engagement. Brain-to-LLM users showed higher memory recall and activation of occipito-parietal and prefrontal areas, similar to the search engine group. Read that carefully. The people who had spent three sessions writing unaided and were then handed the chatbot did not show the depressed pattern. They recalled more, and their measured activation looked like the moderate condition rather than the depleted one.
This is routinely mis-summarized in both directions. It is not true that everything bounces back once the tool is removed, because the withdrawal group did not bounce back within a single session. It is also not true that the tool depresses engagement wherever it appears, because the group that received it last did not show that. What the pattern actually suggests is that order matters: doing the cognitive work first and then using the tool looks different from delegating from the start. That is a hypothesis worth testing properly, and eighteen people across two conditions cannot test it.
The authors asked people not to say thisThe MIT Media Lab project page carries an explicit request to journalists and readers not to describe the findings using words like stupid, dumb, brain rot, harm or damage, and not to write that language models make you stop thinking. The authors state that their conclusions should be treated as preliminary. The framing that made the study famous is a framing its own authors publicly disowned.
One more confound deserves mention. Session four reused prompts participants had already encountered, so familiarity and practice effects are tangled up with the tool switch, the same contamination that makes before and after comparisons of a person's own score so hard to read. Without a control group that stayed on the same tool while facing a familiar prompt, there is no way to separate the two.
4 Its Status Today, and the Formal Critique
As of today the MIT Media Lab paper has not been published in a peer reviewed journal. It was posted to arXiv in June 2025 and revised on 31 December 2025. Its arXiv record carries no journal reference, and the project page still describes peer review as under way rather than complete. Anyone who has described it as a study published in a journal has described it wrongly, and that includes a great deal of the coverage.
In December 2025 a group of researchers from the University of Vienna and Dresden University of Technology posted a formal comment on the preprint. Stanković, Hirche, Kollatzsch and Doetsch open by congratulating the original authors, then set out five categories of concern. Their comment is also on arXiv, and its specifics are more useful than any summary of the debate.
Power. An a priori power analysis for the reported design, using a medium effect size and conventional thresholds, indicates that roughly 159 participants would be required. The study had 54, and the crossover analyses rested on 18.
Multiple comparisons. The commenters note that up to a thousand repeated measures analyses were run across electrode pairs, that the criteria for selecting them are not stated, that the false discovery rate level is never specified, and that reported p values appear to be compared against a plain alpha of .05.
Interpretation of connectivity. As above, a relative measure is being read as an absolute one. The commenters also point out that in the beta band the language model condition showed more significant connections than the unaided condition, which cuts against the tidy gradient.
Reporting inconsistencies. The comment identifies a figure describing a group that does not exist in the design, a session four result described in one direction in the results and the opposite direction in the discussion, and 55 participants stated as completing the study against 54 analyzed with no stated exclusion criterion.
The search engine problem. On the quoting task the language model group underperformed, but the comparison between the search engine group and the unaided group was null. If offloading to an external tool accumulates cognitive debt, the search engine group should have shown it. It did not.
A separate 2026 paper in Frontiers in Developmental Psychology by Pereira Campos and Koff raises the generalizability question: the sample consisted of adults aged 18 to 39 with mature prefrontal development, drawn from one metropolitan area, while the population most exposed to these tools is adolescents. Who a reference sample actually represents, and what follows when it represents a narrow slice, is set out under how norms are constructed. That paper is openly available, and it argues the necessary study has simply not been run on the people who matter most.
5 The Two Studies Built on Self Report
The other two pillars of the argument ask people to report on their own minds, which is the weakest available instrument for this particular question. Gerlich's paper appeared in Societies in January 2025. It combined a survey of 666 participants with interviews, measured how frequently people said they used AI tools, scored critical thinking, and analyzed the results with analysis of variance and correlation. The reported association was negative: heavier self reported AI use went with lower critical thinking scores, with cognitive offloading treated as the mediating path. Younger respondents reported more reliance and scored lower.
Two structural facts limit what that can mean. The design is cross sectional, so every measurement was taken at one moment. A negative correlation between AI use and a thinking score is equally consistent with the tools eroding thinking, with people who think less reaching for the tools more, and with a third variable such as time pressure, education or occupation driving both. Nothing in a single time point can separate those. The paper is also a study of what people say they do, and self reported technology use is a notoriously noisy proxy for actual use. The weakness is not confined to reports of behaviour. Pooled across 41 published studies, what a person believes about their own cognitive standing correlates about .33 with what a test records, leaving a prediction interval roughly 55 points wide, and how far a self estimate sits from a measured score sets a ceiling on anything a survey of this shape can establish.
The paper also received a formal correction. In September 2025 Societies published a correction notice stating that a table in the original article had been published in error, duplicating the contents of an earlier table, and supplying the corrected version. The correction states that the scientific conclusions are unaffected. That is a normal and healthy part of publishing, and it is worth knowing about a paper that the Semantic Scholar index records as having been cited more than a thousand times.
Lee and colleagues at Microsoft Research and Carnegie Mellon surveyed 319 knowledge workers about 936 real tasks, and the paper's own title is refreshingly precise about its outcome: self reported reductions in cognitive effort and confidence effects. Their finding is that higher confidence in the tool goes with less reported critical thinking, while higher confidence in oneself goes with more. They also describe critical thinking shifting rather than vanishing, toward verification, integration and stewardship of the output. The publication record makes no claim about measured performance, and neither should anyone citing it. What separates a report about your own thinking from a scored result is the whole subject of what makes a cognitive test accurate.
Perceived effort is not performanceFeeling that a task took less mental work is compatible with doing it equally well, better, or worse. It is a measure of experience, not output. Treating a drop in felt effort as evidence of a drop in ability is the same error as concluding that a fit runner has become weaker because the hill feels easier.
6 Cognitive Offloading Is Old and Well Studied
The behavior at the center of this argument has a name, a literature, and a history that runs back to the invention of writing. Risko and Gilbert defined it in a 2016 review in Trends in Cognitive Sciences as the use of physical action to alter the information processing requirements of a task so as to reduce cognitive demand. Their examples are deliberately mundane: tilting your head to read rotated text, setting a phone reminder instead of holding an appointment in mind. The review asks two questions, what triggers offloading and what its consequences are, and it is the correct starting point for anyone thinking about this seriously.
The list of technologies accused of ruining human minds is long and unbroken. Writing itself was the first: in the Phaedrus, Plato has Socrates complain that written letters will produce forgetfulness in the souls of learners, because they will trust to external marks instead of remembering from within. Printing was said to encourage shallow reading. The pocket calculator was going to end arithmetic. The telephone directory was going to end the memorization of numbers, and it did, and nothing else happened. Satellite navigation is the current predecessor case.
What the history shows is not that the worries were baseless. Some of them were correct in a narrow, verifiable sense. Most literate adults today cannot recite long passages from memory the way an oral culture could, and the popular idea of a mind that stores pages verbatim turns out to be poorly supported once tested. Most of us no longer know phone numbers. Those are real changes in what people practice, and therefore in what they retain. They are not changes in the capacity to learn a passage or to hold a number, which is a different claim requiring different evidence.
The other thing the history shows is that offloading is usually a net gain. Nobody proposes abandoning writing to recover oral memory. The trade is that external storage frees limited internal resources for other work, and the standard psychometric picture of the six cognitive domains makes it clear why that matters: the CHC framework separates the content you have accumulated, Gc, from the reasoning you can perform on novel material, Gf, and from the temporary holding and manipulation of information, Gwm. A tool that offloads storage is not obviously touching the same thing as a tool that offloads reasoning. Most public argument about AI collapses that distinction on the first sentence.
7 The Google Effect and What Happened to It
The founding study of the digital offloading literature failed a high powered replication, and the popular version never noticed. Sparrow, Liu and Wegner published Google Effects on Memory in Science in 2011, across four experiments. The finding that entered public vocabulary came from the first: when faced with difficult general knowledge questions, participants appeared to be primed to think about computers, shown through slower color naming for computer related words in a modified Stroop task. The paper also reported that people expecting future access to information recalled the information itself less well and recalled where to find it better, an idea framed as the internet functioning as transactive memory. How much of a measured score actually rests on stored content is treated on the page covering memory and measured intelligence. The original record is easy to find, and it has been cited relentlessly.
In 2018 Camerer and colleagues published the Social Sciences Replication Project in Nature Human Behavior, a systematic attempt to replicate 21 experimental social science studies published in Nature and Science between 2010 and 2015. The replications were preregistered, the analysis plans were reviewed by the original authors, and the sample sizes averaged around five times the originals. The Sparrow experiment on the accessibility of computer related words was among those that did not replicate. The project is documented in full.
An independent replication reached the same place. Hesselmann, publishing in PeerJ in 2020, tested the effect again and found no interaction between question difficulty and word type, with a Bayes factor of 5.07 favoring the null, meaning the data were about five times more likely under the hypothesis of no effect. The conclusion states that there was no conclusive evidence that the concept of the internet is automatically activated when people face hard trivia questions. That paper is freely readable.
This matters for a reason larger than one Stroop task. The mechanism by which a single memorable study becomes a fact that everybody knows is exactly the mechanism now operating on a 54 person preprint. A striking result in a prestigious venue gets a name, the name travels, and the replication record, which arrives years later and is far less quotable, never catches up.
The confusion that sustains every headline in this area is the collapse of two separate propositions into one sentence. Separating them dissolves most of the argument, and once you have seen the distinction you cannot read this coverage the same way again.
Claim A: practice shifts
Claim B: capacity falls
What it says
Using a tool changes what you rehearse, so it changes what you know and what you are fluent at
Using a tool reduces your underlying ability to reason, remember and learn
Evidence needed
A comparison of task specific knowledge or skill between users and non users
A normed ability measure, taken more than once, with controls for who chooses the tool
Evidence available
Substantial, across writing, arithmetic, navigation and search
None using an intelligence measure
Reversible
Usually, through practice, since it is a matter of rehearsal
Would depend entirely on the mechanism, which is unspecified
The studies support a weak version of Claim A. They do not touch Claim B, and the headlines are all about Claim B. Take the closest thing to a well designed case: Dahmani and Bohbot, publishing in Scientific Reports in 2020, assessed lifetime satellite navigation experience in 50 regular drivers alongside several facets of spatial memory. Cross sectionally, heavier lifetime use went with worse spatial memory when navigating without it. They also retested 13 participants three years later, and greater navigation use in the interval went with a steeper decline in hippocampal dependent spatial memory. Their paper is careful about what it can support.
Notice what that study measures: spatial memory during self guided navigation. It is a specific skill, closely tied to the thing the tool replaced, and the result is exactly what Claim A predicts. It is not a measure of general ability, and the authors do not present it as one. Whether a measured score can be moved at all, in either direction, is a separate question with its own evidence, set out under raising an IQ score.
Thirteen people, three yearsThe longitudinal arm of the satellite navigation study rests on 13 participants. That is the best longitudinal offloading dataset available on a mature technology after two decades of ubiquity, and it is a fraction of the size required to speak with confidence. It is worth holding that number in mind when someone claims the question about AI is already settled after three years.
9 The Strongest Study in the Field, and What It Shows
There is a large randomized controlled trial on generative AI and learning, it is peer reviewed, it is much better than the three studies everyone cites, and its result is more interesting than the panic. Bastani and colleagues published it in the Proceedings of the National Academy of Sciences in 2025, under a title that does the work honestly: generative AI without guardrails can harm learning.
Nearly a thousand students across ninth, tenth and eleventh grade at a large high school in Turkey were randomized across four ninety minute sessions covering roughly fifteen percent of the mathematics curriculum. Three conditions were tested: a control group with course notes and textbooks only, a group with a standard chatbot interface built on GPT-4, and a group with the same model configured with teacher designed safeguards that supplied hints rather than answers, along with problem specific solutions and common mistakes in the prompt.
The results split cleanly. On assisted practice problems, both AI conditions improved substantially, with the plain chatbot raising performance by around 48 percent and the tutored configuration by around 127 percent relative to the control mean. On the unassisted exam afterwards, the plain chatbot group scored about 17 percent below control. The safeguarded tutor group showed essentially no difference from control, a change that was not statistically significant. The full paper is available through PubMed Central.
Three things follow, and all three are missing from the popular account. First, there is a real cost, demonstrated with a design capable of supporting a causal claim, which none of the three famous studies can do. Second, the cost is a property of how the tool is configured, not of the technology, since a differently prompted version of the same model erased it. Third, and most importantly here, the outcome is performance on a mathematics exam covering material that had just been taught. That is an achievement measure, closer to quantitative knowledge, Gq than to reasoning on novel material, and the distance between school attainment and measured ability has a literature of its own, summarized under IQ and academic achievement. It says nothing about a stable ability index. Students who practiced with answers handed to them learned the material less well. That is a finding about study method, and teachers have known its shape for a century.
10 Why IQ Was Never the Outcome These Studies Could Move
An intelligence battery is not a general purpose detector of cognitive harm, and expecting one to register four sessions of essay writing misunderstands what it samples. A modern battery estimates a common factor across a wide spread of tasks, and derives domain indices from clusters within it, which is how published clinical instruments such as the WAIS-5 are built and what the general factor refers to. The CHC framework organizes these as fluid reasoning, Gf, comprehension knowledge, Gc, quantitative reasoning, Gq, visual spatial processing, Gv, working memory, Gwm, and processing speed, Gs. ACIS administers 20 subtests across those six domains, with subtest scaled scores on a mean of 10 and standard deviation of 3, and composites on a mean of 100 and standard deviation of 15.
Consider what would have to happen for AI use to move a Gf index. Fluid reasoning tasks present novel material with no prerequisite knowledge, and ask for a relation to be inferred under time pressure. There is no content to offload, because the material is unfamiliar by construction. A person who has never written an essay unaided still faces the same matrix or the same series. If chatbot use lowered Gf, the mechanism would have to be something like a general reduction in effortful engagement carried across contexts, which is a strong claim requiring strong evidence, and nobody has produced any.
The domains where a plausible mechanism exists are different ones. Gc is accumulated verbal knowledge, and it is built by reading, writing and retrieval practice, so a genuine long term substitution of composition by generation could in principle show up there, over years rather than weeks. Gwm might respond to habitual delegation of holding and manipulating information. Both are testable predictions. Neither has been tested.
.958
Loading of the ACIS Full Scale IQ composite on the general factor in the technical analysis set.
.9886
Composite omega reliability for the ACIS Full Scale IQ.
N = 2,750
Complete records in the analysis set behind the confirmatory factor and reliability tables.
Those figures come from the published technical manual, and they carry the same qualifications every time they appear. The sample is self selected rather than census based, administration is unsupervised, and the adult reference frame of 3,243 English speaking records aged 16 to 90 is a modeled frame rather than a stratified national sample. Good reliability tells you the instrument gives a consistent reading. It does not license using that reading to detect a cause nobody has demonstrated.
11 The Population Trend Started Before the Technology
Measured cognitive scores in several wealthy countries have been falling for decades, and the decline began with people born before the personal computer. This is the argument most often deployed as proof that technology is degrading minds, and it is the one that most cleanly refutes the AI version of the claim, because the timing does not work.
The Flynn effect describes the substantial rise in measured scores through most of the twentieth century, covered in detail on the dedicated page. It is also why publishers must periodically rebuild their reference frames, a process described under how norms are constructed. Its reversal is documented most rigorously by Bratsberg and Rogeberg, publishing in PNAS in 2018. They used Norwegian military conscription records covering male birth cohorts from 1962 to 1991, with 736,808 valid ability scores from a test comprising arithmetic, word similarities and figural reasoning subtests administered at ages 18 and 19.
Scores peaked with the 1975 birth cohort and fell afterwards at roughly 0.23 points per year across families. The decisive part of the analysis is the within family comparison: because the records identify siblings, the authors could compare brothers raised in the same household, and the within family estimates were larger, in the region of 0.33 to 0.34 points per year. A trend visible between brothers cannot be explained by changing composition of the parent population, by immigration, or by genetic selection. The full paper lays out the reasoning.
Now line the dates up. The 1975 birth cohort was tested around 1993 or 1994. The decline is measured in men who took the test through the 1990s and 2000s. Generative language models reached the public at the end of 2022. Anyone attributing this decline to AI is proposing a cause that postdates its effect by roughly thirty years. The same objection applies with less force but real force to smartphones and social media, which arrive after the turning point too.
What a real population study looks likeNote the scale that was required to establish the reversal credibly: three quarters of a million tested individuals, thirty birth cohorts, a consistent instrument, and a sibling design to rule out compositional explanations. Set that against 54 people writing essays for four months. The gap is not a matter of degree. It is the difference between a dataset that can answer a population question and one that cannot.
12 The Study That Would Actually Answer This
The design required to settle whether AI use changes cognitive ability is well understood, entirely feasible, and has not been carried out by anyone. Describing it is the most useful thing an article on this topic can do, because it converts a shouting match into a specification.
Requirement
Why it is necessary
Status in the current literature
A normed ability measure administered at two or more time points
Change in an individual or a group can only be seen against a stable metric with known error
Absent. No study in this literature administers one
An interval measured in years
Any plausible mechanism for accumulated knowledge or habit operates slowly
Absent. The longest AI study runs four months
Equated or alternate forms
Retesting on identical items measures memory for items, not ability
Not applicable, since no retesting occurs
Controls for self selection into heavy use
People who delegate heavily may differ in ability, motivation and occupation before they ever open a chatbot
Absent. Every AI study here is cross sectional or short and small
Objective usage logs rather than self report
Self reported technology use correlates weakly with measured use
Absent in the survey studies
Preregistration and adequate power
Prevents the analytic flexibility that inflates effects in small samples
Absent in the preprint, present in the randomized education trial
Self selection is the requirement most often waved away and the one most likely to produce a spurious result. If heavy AI users score lower on a reasoning measure, the possibilities include the tools having lowered their scores, lower scoring people having found the tools more useful, and both groups differing on a variable such as job type, time pressure or education that drives use and score independently. A cross sectional survey cannot rank those. A study with a baseline measurement taken before the exposure can, which is precisely why baseline designs are the standard in every other field that studies whether an exposure changes a measured trait. The worked example sits next door: a New Zealand birth cohort measured intelligence at 13, when only seven of its 1,037 members had tried cannabis, and four twin studies then compared siblings discordant for use so that family background could not do the explaining, which is the full ladder of designs that question climbed. The evidence any measurement has to carry before it can support a claim about change is set out under reliability and validity.
Until someone runs it, the correct summary is that the question is open. Not answered in the negative, which would be its own overclaim, but open. The evidence that exists supports a narrow conclusion about practice and study method, and does not reach the broad conclusion about capacity that the coverage announces.
A single test session cannot tell you whether AI use has changed your cognition, and neither can two sessions, and it would be dishonest to suggest otherwise on a page hosted by a test publisher. The obstacle is measurement error, and it is not a small one.
Every score carries a standard error. For the ACIS Full Scale IQ that error is about 1.60 points, which is good, and which reflects a composite omega of .9886 across 20 subtests. But a difference between two scores carries the error of both measurements, so the uncertainty around a change is necessarily wider than the uncertainty around either score on its own. Any real effect of the kind under discussion, if it exists at all, would be small. Small effects buried under compounded measurement error are exactly what individual testing cannot detect, which is why group designs with hundreds or thousands of participants exist. The relationship between error and interpretation is set out on the pages covering the standard deviation of 15 and reliability and validity.
Retesting adds its own problems. Practice effects raise scores on a second administration independently of any change in ability, particularly on timed subtests. Regression to the mean pulls extreme first scores toward the average on retest. Motivation, sleep and mood move scores within a person across days. None of that is a defect of the instrument, and all of it is why claims about raising or lowering your own score require careful design rather than a before and after.
What a full profile can do is describe your current standing across the six domains, placing each index against the percentile distribution and showing whether memory and working memory sit where verbal comprehension does, and whether processing speed is in line with reasoning. That is a useful description of now. It is not a causal history, and a low score does not identify a culprit. ACIS is unsupervised, its normative sample is self selected rather than census based, and it is not a clinical or diagnostic instrument. What an unsupervised online instrument can and cannot establish is treated directly on the page about the accuracy of online IQ tests. It cannot be used to diagnose, to make hiring decisions, to support accommodations, or to establish that any particular habit did or did not affect you.
That boundary is not a house preference. The Standards for Educational and Psychological Testing (2014), issued jointly by the American Educational Research Association, the American Psychological Association and the National Council on Measurement in Education, hold that validity attaches to a specific interpretation for a specific use rather than to a test in the abstract, and that publishers must state the intended uses and the evidence supporting them. APA testing standards and the International Test Commission guidelines on test use make the same demand. A score that is reliable for describing current standing has not thereby been validated for attributing that standing to a cause, and the claim that generative AI lowers intelligence remains, on today's published evidence, untested.
14 Frequently Asked Questions
Does AI make you dumber?
No published study supports that claim, because none has measured intelligence before and after AI use. The strongest evidence, a randomized trial in high school mathematics, shows reduced learning when a chatbot supplies answers, and no reduction when the same model gives hints instead.
Did the MIT ChatGPT study measure IQ?
It did not. Its outcomes were EEG connectivity during essay writing, the ability to quote a line from one's own essay, self reported ownership of the text, and essay scores. No ability index of any kind was administered.
Has the MIT study been peer reviewed?
Not as of today. It was posted to arXiv in June 2025 and last revised on 31 December 2025, and its record carries no journal reference. A formal comment on its methods was posted in December 2025, before the paper itself had completed review.
What does cognitive debt mean in that paper?
It is the authors' framing for an accumulated cost of delegating thinking to a tool. It is a metaphor the paper introduces rather than an established construct with a validated measure, and the formal comment argues the study's measures do not support it.
How many people took part in the MIT study?
Fifty four across the first three sessions, eighteen per condition. The fourth session, which produced the most discussed result, was completed by eighteen people in total, roughly nine in each of the two crossover groups.
What happened in the fourth session?
Participants swapped conditions. Those moving from the chatbot to unaided writing showed reduced alpha and beta connectivity. Those moving the other way showed higher memory recall and activation resembling the search engine group, the opposite of what a simple harm story predicts.
Is cognitive offloading harmful?
Usually not, and often the reverse. Risko and Gilbert's 2016 review treats it as a normal strategy for reducing cognitive demand, and every widely adopted case, from writing to calculators, has been kept because the gains outweigh the narrow skill that was traded away.
Did the Google effect on memory replicate?
The best known experiment did not. The Social Sciences Replication Project reported a failure in 2018, and an independent replication in PeerJ in 2020 found a Bayes factor favoring the null. The original 2011 paper is still cited as settled fact in most popular writing.
What did Gerlich 2025 actually measure?
Self reported frequency of AI tool use against a critical thinking score in 666 people, at a single point in time, supplemented by interviews. The association was negative, but a cross sectional design cannot establish which variable came first.
Was there a correction to the Gerlich paper?
Yes. Societies published a correction in September 2025 noting that a table in the original article had been printed in error, duplicating an earlier table, and supplying the corrected version. The notice states that the conclusions are unaffected.
What did the Microsoft CHI study find?
That knowledge workers with higher confidence in the tool reported less critical thinking, while those with higher confidence in themselves reported more, and that critical thinking shifted toward verifying and integrating output. Every outcome was self reported perception, not measured performance.
Does using a calculator reduce mathematical ability?
The evidence supports a narrow version of the concern: unpracticed arithmetic becomes less fluent, as unpracticed skills generally do. There is no accepted evidence that calculator use reduces quantitative reasoning capacity, and mathematics education has absorbed the tool without such an effect appearing.
Does satellite navigation damage spatial memory?
Dahmani and Bohbot reported in 2020 that heavier lifetime use went with worse unaided spatial memory in 50 drivers, and a steeper three year decline in a follow up group of 13. That is a narrow spatial finding, not a general ability finding.
Is there a randomized trial of generative AI and learning?
Yes. Bastani and colleagues randomized nearly a thousand Turkish high school students across four sessions in a trial published in PNAS in 2025. It is the strongest causal evidence in the field and measures mathematics performance rather than cognitive ability.
Why did the safeguarded version of the AI tutor not harm learning?
Because it was prompted to give hints and withhold answers, so students still had to do the reasoning step. That one model produced harm in one configuration and none in another places the effect in the design of the interaction, not the technology.
Could AI use lower fluid reasoning?
There is no mechanism on offer and no data. Fluid reasoning tasks use novel material with no prerequisite knowledge, so there is nothing about them to offload. A general reduction in effortful engagement carried across contexts is conceivable, but it has never been demonstrated.
Did AI cause the reversal of the Flynn effect?
It cannot have. Norwegian conscription data show scores peaking with the 1975 birth cohort, meaning the decline is visible in men tested from the mid 1990s onward. Generative language models became publicly available at the end of 2022, roughly three decades later.
Why does self selection matter so much here?
Because heavy AI users may differ in ability, occupation and time pressure before they ever start. Without a measurement taken before the exposure, a correlation between heavy use and lower scores fits the tools causing the gap as well as the gap causing the use.
Would taking an IQ test show whether AI has affected me?
No. A single administration describes your standing now and carries no information about change. Comparing two administrations compounds the measurement error of both, adds practice effects and regression to the mean, and still could not isolate one habit as the cause.
Should students avoid AI tools entirely?
The randomized evidence supports attention to design, not a ban. Handing students finished answers during practice reduced later unassisted performance. The same model configured to withhold answers did not, which points at how the tool is used rather than whether it is present.
Can ACIS tell me whether AI has changed my cognition?
It cannot, and no unsupervised online instrument can. ACIS profiles six cognitive domains against a self selected, modeled adult reference frame, which describes current standing only. It is not a clinical or diagnostic instrument and cannot attribute a score to any cause.
Take the assessment
You get a profile, not a number
ACIS measures six CHC domains across 20 subtests and reports each one with its own normed score and confidence interval, so you can see where you are strong and where you are not.