ChatGPT IQ: what happens when a language model sits an IQ test, and what the number means
Since 2023 chatbots have been given Wechsler subtests, Mensa puzzles and bar exams, and the numbers travel: a Verbal IQ of 155, a Mensa chart in the 150s, a bar exam in the top 10 percent. Each number is real, and each is a comparison with a different group of people on a different scale, or with no group at all. This page sets out what each administration measured, and what a human score means beside a machine score.
The name on the logo is a product, not a person: every IQ figure attached to it is a score on one instrument, on one date, with no norm group behind it.
0 Quick Answer
A language model has no IQ in the sense the letters carry for a person, because an IQ is a position within a norm sample of people of the same age on a test whose items the examinee has never seen, and a model has no age, belongs to no norm sample, and may have met the items in training; what a model can have is a score on one instrument under one administration, and every published ChatGPT IQ is a number of that second kind. The best known of them is 155. In March 2023 Eka Roivainen, an assessment psychologist at Oulu University Hospital in Finland, gave ChatGPT five of the six verbal subtests of the WAIS-III and reported in Scientific American an estimated Verbal IQ of 155, above 99.9 percent of the 2,450 Americans in the test's standardization sample. Digit Span could not be given, none of the five nonverbal subtests was given, and the same chatbot could not say who the father of Sebastian's children was. The number describes performance on five verbal tasks against norms drawn from people, and nothing else.
The other numbers in circulation are comparisons with other groups. The GPT-4 technical report describes a simulated bar exam passed with a score around the top 10 percent of test takers, a rank among law graduates on an achievement test. The weekly charts at TrackingAI, read on September 19, 2026, showed the best models in the low 150s on the public Mensa Norway test and in the low to mid 130s on a private test that has never been posted online, with weaker models in the 70s to 100s; both figures are linear rescalings of a raw count of correct answers, and the top of the public chart is the ceiling of the test. The benchmark built for exactly this question, the ARC series described in François Chollet's On the Measure of Intelligence, reports no IQ at all but a share of novel tasks solved, and on September 19, 2026 its leaderboard listed the best system at 95.0 percent on ARC-AGI-2, against a human panel at 100 percent, and the best standard result on the interactive ARC-AGI-3 at about 63 percent.
What the numbers share is that none is on the human scale. A human IQ of 130 is a statement about one person's standing among people of their age with a stated error; a machine 130 on the same items is a statement about a text prediction system on a particular day, with no norm group, no age band and no error of the human kind. Whether using these tools changes a person's own score is a separate question that the page on IQ and artificial intelligence covers; this page is about the practice of scoring the machine.
155
The estimated Verbal IQ from five WAIS-III verbal subtests in Roivainen's 2023 administration, above 99.9 percent of the 2,450 person American standardization sample, with Digit Span not administrable and no nonverbal subtest given.
151
The IQ that the TrackingAI scoring script assigns to a perfect 35 of 35 on the Mensa Norway test, the ceiling of that chart, by our arithmetic on the site's published formula read on September 19, 2026.
0 percent
The score of pure language models on ARC-AGI-2 at its launch on March 24, 2025, on tasks every one of which had been solved by at least two humans in under two attempts.
95.0 percent
The best listed ARC-AGI-2 score on September 19, 2026, for a system listed as GPT-6 Astra at maximum reasoning, beside a human panel listed at 100 percent.
1 What an IQ Is, and What a Model Would Need to Have One
An IQ is not a quantity of intelligence but a rank: the place a person's performance occupies among people of the same age in a reference sample, expressed on a scale with a mean of 100 and a standard deviation of 15, and each part of that definition is something a language model lacks. The construct behind the number goes back to Charles Spearman's observation in 1904 that performance on unrelated mental tasks correlates positively, so that a general factor runs through all of them. A modern battery samples that factor through several domains, verbal knowledge, fluid reasoning, quantitative reasoning, visual spatial ability, working memory and processing speed, as the page on what IQ measures sets out. The American Psychological Association's task force report, Neisser and colleagues' Intelligence: Knowns and Unknowns, summarized in 1996 what such a score does and does not capture in a person.
The number itself is produced by a norm table. A raw score on each subtest is looked up in the distribution of raw scores among people of the same age band in the standardization sample, converted to a scaled score, summed with the others and converted again to a standard score, a process the page on how IQ scores are normed describes. The standard deviation of 15 is the spread of that sample, as the page on the standard deviation of 15 explains, and a percentile is the share of that sample scoring below a given point, as the page on score versus percentile shows.
A model brings none of the four things a person brings to that table. It has no age, so there is no age band to look it up in. It belongs to no norm sample, so a percentile against the sample compares a machine with people who are not its peers on any dimension. It has no standard error of the human kind: a person's score carries an error derived from the reliability of the test across people and occasions, which the page on reliability and validity explains, while a model's variation across runs is a property of its sampling settings rather than of a stable trait. And it may have seen the items: the standardization sample answered questions it had never met, while a system trained on a large share of the written internet may have met discussions, paraphrases or copies of the items. Norms also date. Trahan and colleagues' meta-analysis put the rise in raw performance at 2.31 points per decade, the reason the page on the Flynn effect warns against old tables; a test published in 1997 flatters a human tested in 2023 by roughly six points, by our arithmetic on that figure, and for a machine the correction has no meaning, because it was not born in any year.
2 The First Documented Administration: Roivainen and the WAIS-III Verbal Subtests
The best known ChatGPT IQ, 155, came from a clinical psychologist who gave the chatbot five of the six verbal subtests of the WAIS-III in early 2023, and the number is exactly as meaningful as that description allows. Roivainen's account appeared in Scientific American on March 28, 2023 and in the July 2023 issue. The WAIS-III builds its Verbal IQ from six subtests and its Performance IQ from five, and Roivainen administered Vocabulary, Similarities, Comprehension, Information and Arithmetic. Vocabulary asks for definitions of words, Similarities asks how two things are alike, Comprehension asks why social conventions exist, Information asks general knowledge questions, and Arithmetic asks word problems to be solved without paper. ACIS uses the same task families, and the page on the Vocabulary subtest shows what a typed definition is scored on in a person.
The sixth verbal subtest could not be given. Digit Span asks the examinee to repeat sequences of digits forwards and then backwards, and Roivainen wrote that it cannot be administered to the chatbot, given its lack of the relevant neural circuitry that briefly stores information: a chatbot receives the whole sequence as text and holds nothing over time, so the task measures nothing in it. The page on Digit Span explains why the task is a working memory measure for a person, and the page on working memory tests sets out what the domain contains. The Verbal IQ of 155 was therefore an estimate from five subtests, placed against the 2,450 Americans of the WAIS-III standardization sample, and it exceeded 99.9 percent of them.
Two things were missing from the administration, and Roivainen was explicit about both. The first is the nonverbal half of the test: none of the five performance subtests, which include matrix reasoning, block design and speeded symbol coding, was given, so no Full Scale IQ exists, and the five subtests that were given lean on stored verbal knowledge, the crystallized side of the distinction the page on fluid versus crystallized intelligence describes. The second is reasoning. Asked for the first name of the father of Sebastian's children, the chatbot replied that it lacked sufficient context, and Roivainen's summary was that it fails to reason logically and tries to rely on its vast database. A person scoring 155 on those five subtests would answer that riddle without effort, because in a person the subtests and the riddle draw on the same general factor. In the machine they evidently did not, which is the first sign that the number does not carry across.
3 Exams Are Not IQ Tests: The GPT-4 Report and the Bar Exam
The other figure that circulates as a machine IQ, a bar exam passed in the top 10 percent of test takers, is a rank among a self-selected group of law graduates on an achievement test, and it has no conversion to the IQ scale. OpenAI's GPT-4 technical report, posted to arXiv in March 2023, states in its abstract that the model is less capable than humans in many real world scenarios, and that it exhibits human level performance on various professional and academic benchmarks, including passing a simulated bar exam with a score around the top 10 percent of test takers. Both halves of that sentence matter, and only the second one travelled.
A percentile on a bar exam is a position among the people who sat the bar exam, who have completed law school and selected themselves into the profession. It is not a position among adults of an age, and there is no table that converts one into the other. The page on IQ and academic achievement explains why examinations and ability tests are different instruments even in humans, where they correlate strongly: Deary, Strand, Smith and Fernandes found a latent correlation of 0.81 between general ability at eleven and examination results at sixteen in more than 70,000 English pupils, but that correlation is a fact about people, in whom the same general factor drives both, and it licenses no inference from a machine's exam rank to a machine's ability.
The exam also rewards exactly what a model trained on legal text stores. Bar examinations test doctrine, rules and their application to fact patterns of familiar types, and the written record of that material, including past questions and model answers, is large and public. The report's authors examined the overlap between their evaluations and the training data by matching substrings and reported results with and without the affected items, which is more than most citations of the number mention. What the result shows is that a system with a very large store of legal text can produce passing legal answers. That is a real and consequential capability, and it is a capability of the kind the Information subtest samples in a person, not of the kind a matrix task samples, and the distinction between the two is the whole subject of the next three sections.
4 Matrices in Text: Webb, Holyoak and Lu
The strongest published claim that a language model reasons at a human level on an IQ style task came from a text version of matrix problems built on the rules of Raven's Progressive Matrices, and the comparison it made was with a sample of people, not with a norm table. Taylor Webb, Keith Holyoak and Hongjing Lu reported in Nature Human Behaviour in 2023 a direct comparison between human participants and the text-davinci-003 variant of GPT-3 on a range of analogy tasks. The set included a non-visual matrix reasoning task based on the rule structure of Raven's Standard Progressive Matrices, in which the cells of a matrix were rendered as digits governed by rules such as progression and distribution, together with letter string analogies, four term verbal analogies and story analogies. The authors wrote that GPT-3 displayed a surprisingly strong capacity for abstract pattern induction, matching or even surpassing human capabilities in most settings, and that preliminary tests of GPT-4 indicated even better performance.
The design has a strength that the Wechsler administration lacked. The digit matrices were generated for the study, so they did not exist in any training corpus, and the rule structure is the one that makes matrix items the purest measure of fluid reasoning in a human battery, as the page on matrix reasoning explains. The page on Raven's 2 describes the human instrument the rules come from, and the page on nonverbal IQ tests explains why figural formats are used to reduce the influence of language and schooling. The digit format keeps the rules and discards the figure, which is what a text model needs.
What the result is not is a score. The comparison was between the accuracy of the model and the accuracy of a group of participants on the same problems, and human level in that sentence means the mean of that group, not the mean of a norm sample of a stated age and size. There is no scaled score, no percentile and no interval, and the human group was recruited for the experiment rather than drawn to represent a population. The finding is that on newly generated matrix problems presented as digits, a 2022 language model answered about as often as the people it was compared with. That is an important result about pattern induction in text, and it stands or falls on whether the same ability survives when the surface of the problem changes, which is the question a later study put to it directly.
5 A Cognitive Psychology Battery: Binz and Schulz
When GPT-3 was run through a battery of canonical experiments from cognitive psychology, the profile was jagged in a way no single number can carry: at or above the human level on vignette tasks and a bandit task, and a failure on causal reasoning and on exploration. Marcel Binz and Eric Schulz published Using cognitive psychology to understand GPT-3 in PNAS in 2023, assessing the model's decision making, information search, deliberation and causal reasoning on experiments drawn from the literature. Their abstract lists the results on both sides. GPT-3 solved vignette based tasks similarly to or better than human subjects, made decent decisions from descriptions, outperformed humans in a multi-armed bandit task and showed signatures of model based reinforcement learning. Small perturbations to the vignettes led it vastly astray, it showed no signatures of directed exploration, and it failed, in the authors' word, miserably in a causal reasoning task.
The pattern matters for the practice of scoring machines because a single IQ presumes something that this profile denies. In people, performance across domains correlates positively, which is Spearman's finding and the reason a Full Scale score is a meaningful summary: the page on the six cognitive domains shows how ACIS builds six indices and then a composite that is the best single estimate of what the tasks share. A person far above average on vignette problems is very unlikely to be at the floor on causal reasoning, because the two draw on the same factor. For the model, no such positive manifold across tasks has been demonstrated, and Binz and Schulz's battery is evidence against assuming one. A summary number over a jagged profile is not a summary; it is an average of unlike things.
The second finding, sensitivity to perturbation, is the experimental form of an old psychometric rule. A test item measures ability only if performance on it does not depend on the accidental wording of the item, and human batteries are built by discarding items whose difficulty shifts with irrelevant changes. When a paraphrase of a vignette moves a model from human level to failure, the original success was measuring something about the match between the wording and what the model had stored. That observation, made on a 2022 model, is the same one the counterfactual study made two years later on the analogy tasks, and the same one the ARC Prize team made in 2026 about frontier reasoning systems, and it is the central reason a machine score on a public test is not a measurement of the kind a human score is.
6 The Counterfactual Test: Lewis and Mitchell
The test that separates reasoning from recall is to change the surface of a problem while keeping its logic, and when that was done to the analogy problems on which language models had matched humans, the humans held and the models fell. Martha Lewis and Melanie Mitchell posted Using counterfactual tasks to evaluate the generality of analogical reasoning in large language models to arXiv in 2024. They took the letter string analogy problems used in the 2023 study, created counterfactual variants that test the same abstract reasoning abilities but are likely dissimilar from any pre-training data, and tested humans and three GPT models on both sets. Their abstract reports that while the performance of humans remained high for all the problems, the performance of the GPT models declined sharply on the counterfactual set, and concludes that, despite the earlier successes, the models lack the robustness and generality of human analogy making.
The result does not overturn the 2023 finding so much as locate it. On the original problems, the models matched the people; on problems of the same logical form built from materials the models had not seen, they did not. The difference between the two conditions is the share of the original score that was familiarity rather than reasoning, and for the models that share was large while for the humans it was close to zero. In human testing this is the reason item security exists: the page on IQ test questions explains why a test whose items are public stops measuring ability and starts measuring exposure, and the page on retake control describes how ACIS limits the practice effect that a repeated administration would produce in a person. For a model trained on the public internet, every public item is a potential prior exposure, and a public test is a retake of unknown depth.
The ARC Prize Foundation's technical report for its 2025 competition, posted in January 2026, made the same point about a later generation of systems, describing frontier reasoning performance as fundamentally constrained to knowledge coverage and warning of new forms of benchmark contamination and of knowledge dependent overfitting. The through line from 2023 to 2026 is consistent: scores on public material rise fast, scores on material built to be novel rise more slowly, and the gap between the two is the part of a machine IQ that a human IQ does not contain. That gap is also the design rationale of the private test in the next section.
7 The Weekly Leaderboards: TrackingAI and Its Two Tests
The site most often quoted for a current ChatGPT IQ gives 21 text models and 11 vision models two tests every week, one public and one private, and converts a raw count of correct answers into an IQ with a formula of its own, which is a different object from a score looked up in a publisher's norm table.TrackingAI, run by Maxim Lott, states on its page, read on September 19, 2026, that it quizzes 21 verbal and 11 vision AIs every week. The public instrument is the Mensa Norway test, a 35 item figural matrix test that, for a human taker, has a 25 minute limit and reports scores between 85 and 145, as the page on the Mensa Norway IQ test sets out; the site verbalizes the items for text models and shows the image to vision models. The private instrument is described on the site as a test made by a Mensa member that has never been on the public internet and is in no AI training data; it has 16 items and its questions are withheld from the site's searchable database. If a model refuses an item, the site asks ten times, then uses the most recent answer given and records the refusal.
On September 19, 2026 the charts showed the best models in the low 150s on the Mensa Norway test and in the low to mid 130s on the offline test, with weaker models in the 70s to 100s. The conversion behind those numbers is in the site's own scoring script, read the same day: a raw count on the Mensa Norway test is turned into an IQ by a linear formula, so that a perfect 35 of 35 becomes 151 and each item is worth three points, and the 16 item offline test is rescaled to the same range, so that a perfect score becomes 150, by our arithmetic on the published formula. The site labels the ceiling of each test on its charts. A model shown at 151 has answered every public item, and the number says nothing about how far above the ceiling it would score on a longer test, just as a human 145 on that test is the edge of the scale rather than a measurement of the tail.
The site's public log starts on May 9, 2024, when the model then labelled ChatGPT-4 answered 11 of the 35 public items, which the same formula maps to 79, by our arithmetic; twenty eight months later the top of the chart sits at the ceiling. That is a trajectory of raw counts on a fixed item set, with no one to place a machine against at any point on it, which is precisely what a norm referenced score is not. The gap between the public and the private scores is the site's own contamination check and the reason the private test exists. The page on Mensa's admission test and the page on free versus validated tests set out what a human score on an unsupervised online instrument does and does not establish, and every limit there applies with more force to a machine.
Between a model's raw count of correct answers and a human IQ stand five obstacles, and each one alone is enough to stop the number from meaning what it means for a person. The first is contamination. A model trained on the public internet may have met a public item, a paraphrase of it or a discussion of its answer, and the counterfactual study showed what happens when that possibility is removed: the score falls. A human who has memorized the answer key is not measured by the test. The second is prompt sensitivity. Binz and Schulz found that small perturbations to a vignette could lead GPT-3 vastly astray, and a score that moves with the wording of the question is a score on the wording.
The third is variance across runs. TrackingAI readministers its tests every week and plots a daily moving average because the same model does not produce the same count on the same items from one day to the next. In a person the equivalent quantity is the standard error of measurement, estimated from the reliability of the test across many people and printed as the interval beside a score. A model's run to run spread is a property of its sampling settings and of changes made by its provider, not of the stability of a trait in an individual, and no interval of the human kind can be attached to it.
The fourth is verbal against nonverbal. Roivainen gave verbal subtests only, and a verbal IQ built from vocabulary, general knowledge and word problems is the part of a human profile that reading and schooling build, which is also the part a text corpus supplies. TrackingAI verbalizes a figural test for text models and shows the picture to vision models, and reports the two as separate models; ARC uses grids with no verbal form at all. A human battery samples both sides and reports them as separate indices because they can diverge in a person, and no machine administration has covered both halves with an instrument normed for either. Timed tasks, which the page on processing speed describes, cannot be transferred at all, because a model's speed is a property of hardware and of the provider's settings.
The fifth is the norm problem, and it contains the others. A standard deviation of 15 is the spread of scores among people, a percentile is the share of people below a point, and an age band is a set of people born in the same years; a model is a member of none of those sets. A score that borrows the scale without belonging to the sample borrows the vocabulary of measurement and none of its content. That is not a criticism of the sites and studies that publish such scores, several of which say as much; it is the reason to treat every machine IQ as a raw count with a label, and the reason the ARC benchmarks dropped the label altogether.
9 The Alternative Built for Fluid Intelligence: Chollet's ARC
The benchmark designed for exactly this problem starts from a definition rather than from a test, intelligence as the efficiency with which a system acquires new skills over a scope of tasks given its priors and its experience, and it uses tasks that have no answer key on the internet. François Chollet posted On the Measure of Intelligence to arXiv in November 2019. The paper argues that skill at a specific task is not a measure of intelligence, because unlimited training data or built in priors can buy skill without generalization, and it defines intelligence as skill acquisition efficiency over a scope of tasks, with respect to priors, experience and generalization difficulty. It then introduces the Abstraction and Reasoning Corpus, a set of tasks in which a few input and output grid pairs demonstrate a transformation and the system must produce the output for a new input, built on core knowledge priors such as objects, counting and basic geometry so that a human and a machine start from comparable assumptions.
The design answers the five obstacles of the previous section one at a time. The tasks are unique and novel, so a public answer key does not exist; the evaluation sets used for scoring are kept semi-private or private, so contamination is limited by construction; the score is a share of tasks solved, so no borrowed scale is involved; the format is figural and requires minimal prior knowledge, so the verbal half of the problem is set aside; and the definition builds in efficiency, so that the cost of a solution is part of the result rather than an afterthought. It is, in the terms of a human battery, a fluid reasoning test for machines, built on the same intuition that makes matrix items the core of the fluid side of a human profile, and it reports its results in a unit that makes no claim to be an IQ.
The first version resisted language models for five years. The ARC Prize post on the OpenAI o3 result, published on December 20, 2024, records that GPT-3 scored zero and GPT-4o 5 percent on ARC-AGI-1, and that o3, which the post notes had been trained on the public ARC-AGI-1 training set, scored 75.7 percent on the semi-private evaluation set at a high efficiency setting costing about 26 dollars per task, and 87.5 percent at a low efficiency setting costing about 4,560 dollars per task. Chollet's commentary in the same post is the part to keep beside the numbers: passing ARC-AGI does not equate to achieving AGI, the benchmark is not an acid test for it, and o3 still fails on some very easy tasks, indicating fundamental differences with human intelligence. He also predicted that the second version, then in preparation, would cut the same system to under 30 percent even at high compute, and the launch three months later bore that out.
10 ARC-AGI-2 and ARC-AGI-3: The Leaderboard on September 19, 2026
ARC-AGI-2 was launched on March 24, 2025 with pure language models at zero and the best reasoning systems in single digits, and eighteen months later the leaderboard shows the second benchmark nearly saturated and a third, interactive one on which the best standard result is about 63 percent against a human ceiling of 100. The launch announcement reported that pure language models scored 0 percent, with GPT-4.5 at 0.0 on the semi-private set, that o3-preview at low reasoning scored 4 percent at about 200 dollars per task and o1-pro 1 percent, and that every task had been solved by at least two humans in under two attempts, in live testing of more than 400 people in a controlled setting; the human panel average was 60 percent at about 17 dollars per task, and the grand prize required a score above 85 percent within an efficiency limit. The accompanying paper, ARC-AGI-2: A new challenge for frontier AI reasoning systems, describes the tasks as designed to give a more granular signal at higher levels of fluid intelligence, with the human testing as the baseline.
The ARC Prize 2025 technical report closed the year. Its Kaggle competition, restricted to systems that run within a compute budget, drew 1,455 teams and 15,154 entries, and the top score on the private evaluation set was 24 percent. The report names the defining method of the year as the refinement loop, an iterative per task optimization guided by feedback, notes that four frontier laboratories reported ARC-AGI results in their model cards during 2025, and cautions that frontier reasoning performance remains constrained to knowledge coverage. It also previews ARC-AGI-3, which moves from static grids to interactive environments.
The leaderboard, read on September 19, 2026, showed how far the commercial systems had moved. The leading reasoning systems were listed between 97.5 and 98.5 percent on ARC-AGI-1, level with a human panel listed at 98 percent, and the best of them, a system listed as GPT-6 Astra at maximum reasoning, was listed at 95.0 percent on ARC-AGI-2 at about 1.12 dollars per task, beside a human panel listed at 100 percent at about 17 dollars per task. On ARC-AGI-3, described by the foundation as an interactive benchmark in which agents must explore novel environments, acquire goals on the fly and learn continuously, with every environment human solvable and a full score defined as beating every game as efficiently as a human, the same system's top result under the standard harness was listed at about 63 percent, with a separate group of entries run through a provider adapter listed apart from it. The page on what ARC-AGI is frames the series as a measure of few shot generalization on novel tasks, and none of these figures is an IQ, which is the point: the benchmark that took the norm problem seriously reports a share of tasks and a cost per task, and leaves the human scale to humans.
11 What Each Number Is a Comparison With
Every machine score in circulation is a comparison, and the fastest way to read one is to ask what it is a comparison with, on what scale, and with what ceiling; the table answers those three questions for each documented figure.
Figure
Instrument and administration
Compared with
Scale and ceiling
What it cannot say
Verbal IQ 155 (Roivainen, 2023)
Five of six WAIS-III verbal subtests; Digit Span and the five performance subtests not given
The 2,450 Americans of the WAIS-III standardization sample, via 1990s norm tables
Deviation IQ, mean 100, standard deviation 15, prorated from five subtests
Nothing about fluid reasoning, working memory or speed; a simple riddle failed in the same session
Top 10 percent on a simulated bar exam (GPT-4 report, 2023)
Simulated bar examination
Law graduates who sat the examination
Percentile among exam takers; no IQ conversion exists
Nothing about standing among adults of an age; measures stored legal knowledge
Human level on digit matrices (Webb, Holyoak and Lu, 2023)
New text matrices on Raven's rules plus three analogy tasks, text-davinci-003
The mean accuracy of a recruited human group
Accuracy; no scaled score, percentile or interval
Whether the ability survives a change of surface, tested in 2024
Jagged battery profile (Binz and Schulz, 2023)
Canonical cognitive psychology experiments given to GPT-3
Published human results on each experiment separately
Task by task; no composite
Any single summary number; the tasks did not hang together as they do in people
Sharp decline on counterfactual analogies (Lewis and Mitchell, 2024)
Original and counterfactual letter string analogies, humans and three GPT models
Humans on the same problems
Accuracy on each set
How much of any public score is familiarity rather than reasoning, except that the share is large
Low 150s on the Mensa Norway test (TrackingAI, read September 19, 2026)
35 public figural items, verbalized for text models, shown to vision models, readministered weekly
No norm group; a linear rescaling of the raw count
151 at 35 of 35, the ceiling; the human version reports 85 to 145
Anything above the ceiling; whether the items were in training data
Low to mid 130s on the offline test (TrackingAI, read September 19, 2026)
16 private items written by a Mensa member and never posted online
No norm group; the same kind of rescaling
150 at 16 of 16, the ceiling, by our arithmetic on the site's formula
The reliability of a 16 item instrument, which is unpublished
95.0 percent on ARC-AGI-2 (leaderboard, read September 19, 2026)
Semi-private set of novel grid tasks, best system at maximum reasoning
A human panel listed at 100 percent, at 17 dollars per task
Percentage of tasks solved with cost per task; ceiling 100
Nothing on the IQ scale, by design
About 63 percent on ARC-AGI-3 (leaderboard, read September 19, 2026)
Interactive environments, standard harness, best system
Human solvers, every environment solvable, efficiency relative to humans
Percentage with a human defined ceiling of 100
Nothing on the IQ scale, by design
A human Full Scale IQ on ACIS
20 subtests across six domains, protected administration, retake control
An adult reference frame of 3,243 English speaking records aged 16 to 90
Deviation IQ, mean 100, standard deviation 15, standard error 1.60
Anything about a machine; the comparison runs among people only
The last row is the only one on the human scale, and it is there for contrast rather than for competition. The rows above it are honest numbers with different referents, and the confusion in public discussion comes from placing them in one column. The page on the highest IQ ever recorded makes the same point about human figures quoted from different instruments and decades; a machine 151 belongs to neither conversation.
12 What It Means for a Human Reader
For a person, a score is a comparison with people of their age on a stated scale with a stated error, and the comparison with a model is not on any scale, so the practical question is not whether you are smarter than a chatbot but what your own number is a measurement of. We sell a paid assessment, so treat this section as a disclosure and check it against the rest of the page. ACIS administers 20 subtests across the six domains listed earlier and scores them against an adult reference frame of 3,243 English speaking records aged 16 to 90, with a technical analysis set of 2,750 complete records behind the reliability and factor tables in the technical manual. The Full Scale composite has a reliability of .9886, a general factor loading of .958 and a standard error of 1.60 points, so a reported score is printed with an interval of about three points either side, by our arithmetic on the manual's figures, and every domain index carries its own interval.
The parts of the profile this page has discussed are measured separately. Matrix Reasoning, the fluid task closest to the ARC format and to the digit matrices, has a reliability of .930 and feeds a fluid reasoning index loading at .922 on the general factor; Digit Span, the task Roivainen could not give the chatbot, has a reliability of .890 and feeds a working memory index loading at .788; Vocabulary, which asks for open ended typed definitions rather than multiple choice, has a reliability of .950 and feeds a verbal index loading at .864; and Symbol Search and Coding are timed speed tasks feeding an index loading at .648, the lowest of the six, which is why a speed index is read beside the others rather than as a verdict. The quantitative index loads at .882 and the visual spatial index at .906. The page on what a cognitive test is explains why a battery samples several domains rather than one, and the page on IQ 100 explains what the middle of the scale means for the person standing on it.
The administration is protected, with retake control, which is the human answer to the contamination problem: items are not published, and a second attempt is governed so that a score is not a measure of exposure. The forms are Full Scale with 20 subtests at 50 dollars, Optimized with 13 at 30 dollars and Quick with six at 15 dollars, with a free trial of five subtests that requires no card, 30 days to complete a purchased session and a five day quality guarantee. The public bands run from 90 to 109 Average through 110 to 119 High Average and 120 to 129 Superior to the gifted bands from 130 upward, and each is a band of people. The limits are the ones the page on whether IQ is real sets out for every test: the assessment is self administered and unsupervised, it is not a clinical instrument, and a score describes standing on a defined set of tasks on one occasion. What it does not do is tell you where a machine stands, because no machine is in the frame, and no test built for people can put one there.
13 What the Evidence Supports, Stated Narrowly
Stated as narrowly as the record allows, the practice of giving language models IQ tests supports six sentences, and none of them is that ChatGPT has an IQ of 155. First, on five verbal subtests of the WAIS-III in early 2023, ChatGPT produced answers that, scored against a 1990s human norm sample, gave an estimated Verbal IQ of 155, while the sixth verbal subtest could not be administered, none of the nonverbal subtests was given, and the chatbot failed a simple riddle in the same session. Second, a rank in the top 10 percent of bar exam takers is a rank among law graduates on an achievement test and converts to nothing on the IQ scale. Third, a 2022 language model matched a recruited human group on newly generated text matrices and analogies, and when the same abstract problems were rebuilt from unfamiliar materials in 2024, humans held their level and the models fell sharply. Fourth, on a battery of cognitive psychology experiments, GPT-3 was at or above the human level on some tasks and at the floor on causal reasoning and directed exploration, so that no single number summarizes it. Fifth, the weekly leaderboard scores are linear rescalings of raw counts on a 35 item public test whose ceiling the best models had reached by September 19, 2026, and on a 16 item private test on which they sat lower, with neither number resting on a norm group. Sixth, the benchmarks built for fluid intelligence report a share of novel tasks solved and a cost per task, and on September 19, 2026 they showed the best system at 95.0 percent on ARC-AGI-2 against a human panel at 100 and at about 63 percent on the interactive ARC-AGI-3, whose every environment humans have solved.
What the evidence does not support is any placement of a model on the human scale. It does not support reading a machine 151 as rarer than a human 145, because the first number is a ceiling on a rescaled count and the second is a position in a sample. It does not support reading a verbal estimate as a Full Scale score, or an exam rank as an ability score, or a mean matched in one experiment as an ability that survives a change of surface. And it does not support the reverse claim either: nothing on this page shows that the systems lack ability, only that the instruments built for people cannot measure whatever they have, and that the instruments built for them measure something they still fall short on when it is interactive and novel.
The open question is not the one in the headlines. It is whether a measure of skill acquisition efficiency of the kind Chollet defined can be given a scale, a norm group of systems and an error of its own, so that a machine result becomes a measurement rather than a count. Until it can, the honest summary of a ChatGPT IQ is a raw score with a label borrowed from psychometrics, and the honest summary of a human IQ is the thing the page on what IQ is describes: a position among people, with an error, on a scale that was built for them.
Every figure above is traceable to one of the following, and each is linked at the point where it is used. Web pages are quoted as read on September 19, 2026 and their figures change weekly; the conversion of the TrackingAI raw counts to IQ ceilings, the Flynn correction for a 1997 test, and the interval around an ACIS score are our arithmetic on published figures and are labelled as such where they appear. ACIS reliability, loading and standard error figures are from the technical manual, version 1.4.
Roivainen E. I gave ChatGPT an IQ test. Here's what I discovered. Scientific American, opinion, March 28, 2023, and the July 2023 issue.
Webb T, Holyoak K J and Lu H. Emergent analogical reasoning in large language models. Nature Human Behaviour, 2023, volume 7, issue 9, pages 1526 to 1541.
Lewis M and Mitchell M. Using counterfactual tasks to evaluate the generality of analogical reasoning in large language models. arXiv 2402.08955, 2024.
Chollet F. On the Measure of Intelligence. arXiv 1911.01547, 2019.
Chollet F, Knoop M, Kamradt G, Landers B and Pinkard H. ARC-AGI-2: A new challenge for frontier AI reasoning systems. arXiv 2505.11831, 2025, revised January 2026.
Chollet F, Knoop M, Kamradt G and Landers B. ARC Prize 2025: Technical report. arXiv 2601.10904, January 2026.
ARC Prize Foundation. Announcing ARC-AGI-2 and ARC Prize 2025. arcprize.org, March 24, 2025.
ARC Prize Foundation. OpenAI o3 breakthrough high score on ARC-AGI-Pub. arcprize.org, December 20, 2024.
Lott M. Tracking AI: AI IQ test results and methodology. trackingai.org, read on September 19, 2026, including the site's public scoring script and daily score log.
Spearman C. "General intelligence," objectively determined and measured. The American Journal of Psychology, 1904, volume 15, issue 2, pages 201 to 292.
Neisser U, Boodoo G, Bouchard T J, Boykin A W, Brody N, Ceci S J, Halpern D F, Loehlin J C, Perloff R, Sternberg R J and Urbina S. Intelligence: Knowns and unknowns. American Psychologist, 1996, volume 51, issue 2, pages 77 to 101.
Trahan L H, Stuebing K K, Fletcher J M and Hiscock M. The Flynn effect: A meta-analysis. Psychological Bulletin, 2014, volume 140, issue 5, pages 1332 to 1360.
Deary I J, Strand S, Smith P and Fernandes C. Intelligence and educational achievement. Intelligence, 2007, volume 35, issue 1, pages 13 to 21.
15 Frequently Asked Questions
What is the IQ of ChatGPT?
It has no IQ in the human sense, because an IQ is a position within an age based norm sample of people. The most quoted figure, an estimated Verbal IQ of 155, came from five WAIS-III verbal subtests given by a clinical psychologist in 2023, with no nonverbal subtests and no Digit Span, and it describes that administration only.
Did ChatGPT really score 155 on an IQ test?
Yes, as an estimate from five of the six WAIS-III verbal subtests, Vocabulary, Similarities, Comprehension, Information and Arithmetic, scored against the 2,450 person American standardization sample of a test published in 1997. Digit Span could not be administered, the five performance subtests were not given, and the same chatbot failed a simple riddle.
Can a language model have an IQ at all?
Not on the scale used for people. An IQ requires an age band, a norm sample the examinee belongs to, items the examinee has never seen and a standard error derived from human reliability. A model has no age, belongs to no sample, may have met the items in training, and varies across runs for reasons unrelated to any trait.
Is GPT-4 smarter than humans on IQ tests?
The GPT-4 technical report claims human level performance on professional and academic exams, including a simulated bar exam in the top 10 percent of takers, while stating that the model is less capable than humans in many real world scenarios. An exam rank among law graduates is not an IQ, and no study has put GPT-4 through a normed battery.
What IQ do the best AI models score on the Mensa Norway test?
On September 19, 2026 the TrackingAI charts showed the best models in the low 150s on that 35 item test, which is the ceiling: the site's formula assigns 151 to a perfect score. The human version of the test reports scores only between 85 and 145, and the figure is a rescaled raw count rather than a normed score.
Which AI has the highest IQ?
No AI has an IQ, so the question reduces to which model tops a given chart on a given week, and the answer changes weekly. On the public Mensa Norway test several models sat at the ceiling on September 19, 2026, which cannot separate them, and the private 16 item test on the same site ranked them lower and differently.
Does an AI passing the bar exam mean it has a high IQ?
No. A bar exam percentile is a rank among people who completed law school and sat the exam, on a test of legal doctrine and its application, much of which is public text. It has no conversion to the IQ scale and says nothing about fluid reasoning, working memory or speed, the domains a human battery samples alongside knowledge.
Which WAIS-III subtests were given to ChatGPT, and which could not be?
Roivainen administered Vocabulary, Similarities, Comprehension, Information and Arithmetic, five of the six subtests that make up the WAIS-III Verbal IQ. Digit Span could not be administered because it requires holding a sequence briefly in memory, and none of the five performance subtests, which include matrix reasoning and block design, was given.
Why could Digit Span not be administered to a chatbot?
Digit Span asks the examinee to hear a sequence of digits and repeat it forwards, then backwards, which measures the brief storage and manipulation of information. A chatbot receives the whole sequence as text and holds nothing over time, so, in Roivainen's phrase, it lacks the relevant neural circuitry that briefly stores information, and the task measures nothing in it.
What did Webb, Holyoak and Lu find with the matrix problems?
In 2023 they compared the text-davinci-003 version of GPT-3 with human participants on newly generated text matrices built on the rules of Raven's Progressive Matrices, plus letter string, verbal and story analogies. The model matched or surpassed the human group in most settings, and preliminary GPT-4 tests did better, as a comparison of accuracies rather than a normed score.
What did the counterfactual analogy study show?
Lewis and Mitchell rebuilt the letter string analogies from materials unlikely to appear in training data and tested humans and three GPT models on both versions. Human performance stayed high on both; the models declined sharply on the counterfactual set, which the authors read as a lack of the robustness and generality of human analogy making.
What did the cognitive psychology battery of Binz and Schulz find?
GPT-3 solved vignette based tasks similarly to or better than people, made decent decisions from descriptions, outperformed humans on a multi-armed bandit task and showed signs of model based learning, but small changes to the vignettes led it far astray, it showed no directed exploration, and it failed a causal reasoning task, a profile no single number summarizes.
How does TrackingAI score its IQ tests?
It gives 21 text models and 11 vision models the 35 item Mensa Norway test and a 16 item private test weekly, verbalizing items for text models, and converts each raw count to an IQ with a linear formula of its own: a perfect public score becomes 151 and a perfect private score 150. No norm table is used.
What is ARC-AGI, and how do models score on it?
ARC-AGI is a series of benchmarks built on Chollet's 2019 definition of intelligence as skill acquisition efficiency on novel tasks. On September 19, 2026 the leaderboard listed the best systems near the human panel on ARC-AGI-1, at 95.0 percent on ARC-AGI-2 against a human 100, and at about 63 percent on the interactive ARC-AGI-3, which is fully human solvable.
Why is a public IQ test a problem for testing a model?
Because a model trained on the public internet may have met the items, their paraphrases or their answers, and a score on familiar material measures exposure rather than ability. The counterfactual study showed the effect directly, and the private test on TrackingAI and the semi-private ARC evaluation sets exist to limit it.
Why do model scores vary from one run to the next?
Language models sample their answers, so the same items can produce different counts on different days, and providers change their systems over time. TrackingAI readministers weekly and plots a daily moving average for that reason. The variation is a property of sampling and of updates, not the standard error of a stable trait in an individual.
What is the norm problem in giving a model an IQ test?
An IQ is a position within a sample of people of the same age, the standard deviation of 15 is the spread of that sample, and a percentile is the share below a point. A model belongs to no such sample, so a score on that scale borrows the vocabulary of measurement without the comparison that gives it meaning.
Can I compare my IQ with a ChatGPT score?
Not meaningfully. Your score is a comparison with people of your age on a stated scale with a stated error; a machine score is a raw count on one instrument with a borrowed label, no norm group and no interval. A human 145 and a machine 151 on the same public test are different objects.
How does ACIS measure a person differently from a benchmark?
ACIS scores 20 subtests across six domains against an adult reference frame of 3,243 English speaking records aged 16 to 90, reports a Full Scale IQ with a reliability of .9886 and a standard error of 1.60 points, prints an interval beside every index, and protects its items with retake control. A benchmark reports tasks solved and cost per task.
Does using ChatGPT change my own IQ?
No published study has measured intelligence before and after use of these tools, so the question is open rather than answered. The studies cited for a decline measured brain signals, self reports or a mathematics exam, not a normed ability score, and the documented costs concern how material is learned when answers are supplied rather than worked out.
What would a fair IQ test for a machine look like?
It would use items that cannot be in training data, report a share of tasks solved with the cost of solving them, compare against a human panel that solved every task, and eventually define a norm group of systems and an error of its own. The ARC benchmarks supply the first three elements; nobody has yet supplied the fourth.
Take the assessment
You get a profile, not a number
ACIS measures six CHC domains across 20 subtests and reports each one with its own normed score and confidence interval, so you can see where you are strong and where you are not.