Documented administration

Muhammad Ali's army test was real. The IQ number attached to it was not

In 1964 Muhammad Ali sat an Armed Forces qualifying test and was reported to have scored at roughly the 16th percentile, below the cutoff then in force. That administration is documented. The IQ figures of 78 and 83 that later attached to it are reported, and a percentile is not an IQ point score.

Muhammad Ali standing in a boxing ring in white trunks and red gloves, shouting.
A percentile says where a person ranked. A standard score says how far from a mean they sat. The two scales answer different questions.

0 Quick answer: a real test, the wrong label

Muhammad Ali has no documented IQ score. He has a documented aptitude test result, and the two are not the same thing. In 1964 Ali sat the mental qualifying examination used by the United States armed forces and was reported to have placed at roughly the 16th percentile, below the qualifying floor then in force. He was classified 1-Y, meaning fit for service only in a national emergency. That much is on the record.

The complication is what happened to that result afterwards. An aptitude screen scored in percentiles was gradually retold as an intelligence quotient scored in points. Press accounts settled on 78. Columnist Clarence Page, writing on 8 June 2016, reported a different figure of 83, attributing it to the historian Anthony O. Edmonds and to material in FBI files. Neither number is a published IQ test score, because no IQ test publisher has ever reported administering a test to Ali.

This page treats the case as a worked example rather than a verdict. The Armed Forces Qualification Test is scored as a percentile rank against a reference group. An IQ is a standard score expressed in points around a mean. Turning one into the other requires assumptions about which population, which year and which distribution, and those assumptions are almost never stated when a celebrity number is quoted.

1964

Year of the documented Armed Forces qualifying test administration.

16th percentile

Reported result. A rank among test takers, not a point score.

Zero

Published administrations of any IQ battery to Muhammad Ali.

ACIS has not tested Muhammad AliNo number anywhere on this page is an ACIS measurement, an ACIS estimate, or an ACIS opinion about Muhammad Ali's intelligence. ACIS has never administered any assessment to him and never will, as he died in 2016. Every figure below is attributed to the source that reported it, with an explicit label saying how strong that source is.

If you want the underlying mechanics rather than the biography, the two scales are separated in detail on the percentile and standard score comparison, and the conversion points themselves are tabulated on the IQ percentile chart. This page does not restate them. It shows what happens when they are confused in public for sixty years.

1 What is actually documented about the 1964 test

Strip away the retellings and a short, checkable sequence of events remains. Ali, then still fighting as Cassius Clay, was called for pre induction screening in 1964, the year he took the heavyweight title from Sonny Liston. He sat the written mental qualifying examination and did not reach the qualifying standard. The Selective Service reclassified him 1-Y, a category reserved for registrants acceptable only in time of war or declared national emergency.

Rhiannon Walker, writing for Andscape on 11 January 2017, records the classification plainly: because of unsatisfactory scores on the mental acuity test, he was classified 1-Y in 1964, and in 1966 the military lowered its draft eligibility standards and Local Board No. 47 reclassified him 1-A. Mark Weisenmiller, writing for the History News Network on 4 October 2015, adds that Ali struggled particularly with the mathematical questions.

The later legal history is a matter of published record. Ali refused induction in 1967, was convicted, and the Supreme Court reversed that conviction in Clay v. United States, 403 U.S. 698 (1971). That opinion turns on the handling of his conscientious objector claim rather than on his test score, so it confirms the draft dispute without confirming any number. It is worth reading precisely because people cite it as though it settled the testing question. It did not.

Notice what this record contains and what it does not. It contains a date, an instrument category, a classification outcome, and a policy change. It does not contain a score report, a norm table, an examiner's notes, testing conditions, or a confidence interval. In clinical or research use, a score without those things is not interpretable, which is a point made in every treatment of measurement quality.

Documented, not measuredThat an administration happened is documented. What it yielded, in what units, against which reference group, is reported second hand. Those are two different evidential standards, and celebrity IQ writing routinely collapses them into one.

This is still an unusually strong case by the standards of the genre. Most famous IQ claims rest on nothing at all: a biographer's guess, a fan forum, or a number that appeared in a listicle and then propagated. The pages on the highest recorded IQ claims show how thin that evidence usually is. Ali is one of the very few names where a genuine, dated, official administration sits underneath the story. That is exactly why the case is worth working through carefully instead of dismissing. Sorted by what has to exist before a claim can be checked, the audit of thirty five famous IQ claims finds four with any cognitive administration behind them and twenty five resting on a figure whose origin nobody can name, which is the company this 1964 record keeps.

2 What the Armed Forces Qualification Test is, and is not

The AFQT was built to answer an institutional question, not a psychological one: can this applicant be trained efficiently for military work. The National Longitudinal Surveys programme, which has used AFQT data in research for decades, describes it as a general measure of trainability and a primary criterion of enlistment eligibility for the armed forces. That is a statement about administrative function, and it deliberately avoids the word intelligence.

In its modern form the AFQT score is not a separate test at all. The official ASVAB documentation states that it is a composite derived from four subtests of the wider battery: Arithmetic Reasoning, Mathematics Knowledge, Paragraph Comprehension and Word Knowledge. Two of the four are arithmetic. The other two are reading and vocabulary. Nothing in that composite samples visual spatial ability, working memory span or processing speed.

That narrowness matters for Ali specifically, and section nine returns to it. A composite made of arithmetic and reading, delivered as a timed written paper, is a fair screen for a soldier who will read manuals and calculate. It is a poor sample of general ability, because it leaves several broad domains untouched. In the terms used across the CHC framework, the AFQT loads on Gq quantitative knowledge and Gc comprehension knowledge and largely ignores Gv, Gwm and Gs.

The scoring convention is the second half of the story. The same official documentation states that AFQT scores are reported as percentiles between 1 and 99, and that an AFQT percentile score indicates the percentage of examinees in a reference group that scored at or below that particular score. The output is a rank. It is not a quantity of anything.

PropertyAFQTA standard IQ battery
PurposeEnlistment eligibility and trainability screeningEstimation of general cognitive ability and its subdomains
Score unitPercentile rank, 1 to 99Standard score, mean 100 and standard deviation 15
Content sampledArithmetic, mathematics knowledge, reading, vocabularyMultiple broad domains including reasoning, spatial, memory, speed
Reported errorNot routinely reported to the examineeStandard error of measurement and confidence interval reported
InterpretationAbove or below an administrative cutoffProfile of relative strengths against a norm sample

Researchers do sometimes use AFQT scores as a proxy for general ability, because in large samples the composite correlates respectably with broader measures. That is a defensible research move at the level of populations. It is not a licence to relabel one person's percentile as their IQ. The distinction between what a measure supports about a group and what it supports about an individual is developed further on the page comparing IQ scores with ASVAB and AFQT results. The battery's own owner draws the same line: an executive note from the Office of People Analytics titled Appropriate Use of ASVAB Scores lists only three validated uses, enlistment selection, occupational classification and career exploration, which is why the inventory of instruments in professional use files the ASVAB among aptitude screens rather than among intelligence batteries.

3 A percentile is a rank, an IQ is a distance

This is the sentence the whole Ali story turns on: a percentile tells you how many people you finished ahead of, and a standard score tells you how far you sat from the average. They are different questions, and they produce numbers that look alike and mean nothing alike.

A percentile of 16 means that roughly 16 in every 100 people in the reference group scored at or below that point. It is an ordinal position. Its spacing is not even: the gap in raw ability between the 50th and 55th percentile is much smaller than the gap between the 94th and the 99th, because scores bunch in the middle of a distribution and thin out at the edges.

A standard score of 85 means something structurally different. It means one standard deviation below a mean of 100, on a scale where the standard deviation is fixed at 15. Its spacing is even by construction. Every 15 points is one standard deviation, everywhere on the scale. That property is what makes standard scores addable, comparable across subtests and usable in a profile, and it is set out on the page on the 15 point standard deviation.

Question you are askingScale that answers itWhat the number means
How many people did I finish ahead of?Percentile rankOrdinal position in a named reference group
How far am I from the average?Standard score, mean 100 SD 15Distance from the mean in fixed units
How unusual is this result?Rarity, one in XFrequency of that result or better in the population
Am I above or below a policy threshold?Cutoff decisionA yes or no, set by an institution, not by the test

You can move between the first two scales, but only under conditions. You need the trait to be approximately normally distributed, you need both scales referenced to the same population, and you need to know which year's norms are in play. Meet those conditions and the 16th percentile corresponds to a standard score of about 85, and the 18th to about 86. Those are arithmetic facts about a normal curve, and the percentile chart lists the full correspondence.

They are not facts about Muhammad Ali. The conditions were not met in his case, for reasons the next four sections set out one at a time. Quoting 85 as his IQ would be the same error as quoting 78, dressed in better arithmetic. If you want to see how the same rank converts into a frequency statement, the rarity calculator does that conversion explicitly, and it will show you how quickly small percentile changes stop mattering near the middle of the distribution.

4 Where the number 78 comes from

The figure 78 is the most repeated number in this story and the least anchored. It circulates in press accounts, boxing histories and quiz pages as Ali's army IQ score. What it does not come with, in any account reviewed for this page, is a named instrument, a date of scoring, a scoring organisation, or a reference group. It is reported, and reported widely, but its provenance is not established.

There is also an internal problem with it that is rarely noticed. The reported percentile result is 16. The AFQT is scored in percentiles running from 1 to 99. If 78 were an AFQT percentile it would describe a well above average performance, which contradicts the documented outcome of failing to reach the cutoff. So 78 cannot be the AFQT percentile. It is a number on some other scale.

Which other scale is where the trail goes cold. It could be a raw or converted score on a mid century service instrument, several of which reported on scales unrelated to the modern mean 100 metric. It could be a later conversion performed by a journalist. It could be a transcription of something else entirely. Honest treatments of the case say they do not know, and this page says the same.

Weisenmiller's History News Network account offers a useful counterweight, because it reports a bare score of 16 and notes that when the Pentagon later set the floor at 15, Ali's result sat a single point above it. That is a coherent story told entirely in percentile units, with no IQ figure in it at all. It suggests the original record was a percentile all along and that the point score arrived later, from outside.

Label: reported, provenance unestablishedThe 78 is reported by many secondary sources and traceable to none of them in the form of a score report. It should never be written as "Muhammad Ali's IQ was 78". At most it can be written as "a figure of 78 has been widely reported in press accounts, without an identified instrument".

It is worth being blunt about why this matters beyond one man's reputation. A number with no instrument attached is not a weak measurement, it is not a measurement. The habit of quoting it anyway is what makes celebrity IQ lists worthless as evidence, and it is the reason reliability and validity are treated as prerequisites rather than footnotes. A score you cannot trace cannot be checked, and a score that cannot be checked cannot be wrong, which sounds like a strength and is in fact the defining weakness.

5 The competing figure of 83, and why two numbers exist

A second number circulates alongside the first, and it has a slightly better paper trail. Clarence Page, in a syndicated column dated 8 June 2016, wrote of Ali that he barely earned a certificate of attendance and scored 83 on an IQ test administered by the army. Page attributes the finding to Anthony O. Edmonds's Muhammad Ali: A Biography, published by Greenwood Press in 2006, and to material Edmonds drew from FBI files.

Edmonds is a real and checkable authority. He is a professor of history at Ball State University who has written extensively on the Vietnam era, and the biography exists as a catalogued 144 page volume with bibliographical references and an index. That is a considerably stronger chain than the 78 has: a named historian, a named publisher, a named archival source, relayed by a named columnist on a fixed date.

Stronger is not the same as sufficient. This page has confirmed that the Edmonds volume exists and that Page attributes the figure to it. It has not been able to read the underlying page of the biography or the FBI document behind it, so the correct label is reported by a named secondary source, attributing to a named primary source, unverified at the primary level. Anyone who writes it more confidently than that is overstating what they checked.

The existence of two competing figures is itself informative. In a properly documented assessment there is one score report, and disagreements are about interpretation rather than about what the number was. Where two different point scores circulate for the same sitting, at least one is wrong, and the likelier explanation is that both are downstream reconstructions of a percentile record that never contained a point score at all.

FigureWhat it is presented asNamed sourceEvidential label
16th to 18th percentileResult on the 1964 qualifying testClarence Page, 2016; Weisenmiller, 2015Reported, consistent across sources
78Army IQ scorePress tradition, no identified instrumentReported, provenance unestablished
83Score on an army administered IQ testPage, 2016, citing Edmonds, 2006, citing FBI filesReported, primary source unverified here
About 85Standard score matching the 16th percentileNormal curve arithmeticEstimated, valid only under stated assumptions
Any single "true" IQAli's general cognitive abilityNoneUnsupported

Read that table as the honest answer to the search query. The percentile facts are reasonably solid and agree with each other. The point scores do not agree, do not carry instruments, and should not be repeated as if they did. That pattern, solid ordinal record and invented cardinal gloss, recurs across almost every entry in the celebrity IQ genre.

6 What converting that percentile into an IQ would require

Suppose you wanted to do the conversion properly rather than dismissing it. The list of things you would need is longer than most people expect, and Ali's case supplies almost none of them.

You would need, first, the exact reported statistic and its units, not a second hand recollection of it. You would need the identity of the reference group the percentile was computed against, along with its year, its age range and its composition. You would need evidence that the underlying trait is distributed approximately normally in that group, since the percentile to standard score mapping depends entirely on the shape of the distribution.

You would then need to establish that the reference group is comparable to the population that defines the IQ scale you are converting into. This is the step that quietly fails most often. An IQ of 100 is defined as the mean of a specific normed population. A percentile computed against military applicants, or against a wartime service population, is anchored to a different group with a different mean, and the two anchors do not coincide.

You would need the reliability of the instrument, in order to attach a confidence interval. A single point estimate with no interval is a misleading object regardless of how carefully it was derived. And you would need the testing conditions, because a timed written test taken under conditions that impede reading measures something narrower than the same construct measured under accommodation.

Reference group

Unknown for the 1964 administration in the sources reviewed.

Reliability

Not reported to the examinee, so no confidence interval is possible.

Conditions

Timed written format, with reading difficulty reported.

Every one of those requirements is standard practice, not a purist's objection. They are what norming a test means, and they are the reason a modern score report carries a band rather than a point. When a conversion is offered without them, the missing assumptions have not gone away. They have simply been made silently, and usually generously.

The unstated assumption problemConverting a 1964 aptitude percentile into a modern IQ point score requires assuming that two different reference populations are equivalent, that the trait is normally distributed in both, and that testing conditions were adequate. State those three assumptions out loud and almost no one will accept the conversion. That is why they are never stated.

The same reasoning applies to ACIS itself, and it would be dishonest to apply it only to a 1964 army file. ACIS reports Full Scale IQ with a standard error of measurement of about 1.60 points, and it reports composites on the conventional mean 100 and standard deviation 15 metric against an adult reference frame of 3,243 English speaking records aged 16 to 90. That frame is a modelled adult frame built from self selected volunteers who chose to take an unsupervised online test. It is not a census sample, and every interpretation on the technical manual is bounded by that fact.

7 The reference group problem in Ali's era

A percentile is meaningless without naming the group it is a percentile of, and the group used in 1964 was not the American public. This is the single most underappreciated fact in the whole case, and it undermines every conversion that has ever been attempted on Ali's score.

For decades the armed services interpreted applicant test results against World War Two era data rather than against a contemporary civilian sample. A National Research Council workshop report on enlistment standards, published by the National Academies, marks the change in tone precisely when it observes that no longer do the test scores of today's recruits have to be interpreted with test data from the World War Two era. The relief in that sentence is the point. Until 1980 they did.

So a 1964 applicant's percentile placed him against a mobilisation era reference population, drawn overwhelmingly from men entering service two decades earlier, with the educational profile of the 1940s. Whatever that comparison is, it is not a rank against American adults of 1964, and it is emphatically not a rank against the population that defines a modern IQ scale. The anchor is simply somewhere else. The only two documented presidential scores have the same hole in them: Kennedy's 119 and Nixon's 143 came from an Otis group test whose form and norm year are named nowhere, so the reference population behind two otherwise documented figures is unknown.

The same National Academies account records how fragile these anchors are in practice. It reports that the battery in use since at least 1976 had been misnormed, and that the norming error caused AFQT percentile scores to be inflated. The consequence was that a large number of applicants who did not meet the standard were nonetheless enlisted before the error was found. That is an official acknowledgement that percentile scores can be systematically wrong by enough to change hundreds of thousands of decisions.

The correction came through the Profile of American Youth study in 1980, which administered the battery to a nationally representative sample of young people and gave the scale a defensible civilian anchor for the first time. The current anchor is more recent still: the official documentation describes today's reference group as a sample of 18 to 23 year old youth from a national norming study conducted in 1997.

Three different anchors, one scale nameA percentile from 1964, one from 1978 and one from today all read as a number between 1 and 99, and all three are referenced to different populations. Comparing them directly, or converting any of them into modern IQ points, treats three incompatible anchors as one. The label on the scale stayed the same while the meaning underneath it moved.

Anyone who quotes Ali's 16th percentile as though it were a rank among people alive today is making exactly this mistake. The number is real. The comparison group it refers to is a wartime service population from the 1940s, which is why how a test is normed is not a technical afterthought but the thing that determines whether a score means anything at all.

8 The same score, two opposite verdicts

Nothing demonstrates the difference between a measurement and a decision more cleanly than what happened next: Ali's result did not change, and its verdict reversed completely.

In 1964 the qualifying floor sat at the 30th percentile, as Page's column records, and Ali's reported result fell below it. He was classified 1-Y and was, for practical purposes, out of the draft. In early 1966, with the war in Vietnam escalating and manpower requirements rising, the Department of Defense lowered the mental qualification floor to the 15th percentile. Weisenmiller reports the directive setting a recorded score of 15 as sufficient, leaving Ali's 16 a single point above the new line.

He did not retake anything. He did not become more or less able. A policy threshold moved beneath a fixed result, and a man who had been officially unqualified became officially qualified. Walker's account records the administrative consequence: Local Board No. 47 reclassified him 1-A in 1966, which set in train his refusal of induction, the loss of his titles and years of his athletic prime, and eventually the 1971 Supreme Court reversal.

The wider policy was the manpower programme that lowered accession standards from 1966 onward and brought several hundred thousand men into the services who would previously have been screened out. Whatever one concludes about that programme, its existence proves the general point. A cutoff is an institutional choice about how many people an organisation needs and what risk it will tolerate. It is not a property of the person being measured, and it is not a property of the test.

30th

Qualifying percentile floor in force in 1964, per Page, 2016.

15th

Floor after the early 1966 change, per Page and Weisenmiller.

Unchanged

Ali's underlying result across both verdicts.

The lesson generalises well beyond the draft. Every threshold applied to a test score is a decision layered on top of a measurement: an admissions cutoff, a hiring screen, a society's entry requirement, a diagnostic criterion. The score can be perfectly reliable and the threshold still arbitrary, and moving the threshold changes who passes without changing anyone's ability. That is why interpreting a score always requires asking who set the line and why.

It is also why ACIS reports a band rather than a verdict. A Full Scale IQ from ACIS arrives with a standard error of measurement of about 1.60 points and an explicit statement that the administration is unsupervised, so that a reader can see how much of the number is signal and how much is measurement noise before anyone builds a decision on top of it. A test that hands you a threshold verdict without the interval has made the choice for you and hidden the fact.

9 Reading difficulty and what a written timed test then measures

Ali's difficulty with reading and writing is reported consistently across sources, and it bears directly on what a timed written examination could have measured in his case. It deserves to be stated plainly, without euphemism and without being used as a rescue.

Ali is described as dyslexic, with the resulting difficulties in reading and writing, and his 1964 failure is attributed specifically to writing and spelling skills that fell below the required standard. He struggled through school in Louisville and finished near the bottom of his graduating class. In his later memoir written with his daughter Hana, he addressed the difficulty directly and named it. He also spent years of his public life campaigning on literacy, which is not the behaviour of a man hiding the fact.

Now apply that to the instrument. The AFQT composite is arithmetic and reading, delivered on paper, against the clock. For a person whose access to printed text is impaired, a timed written test does not cleanly measure quantitative reasoning or verbal knowledge. It measures those things filtered through the decoding bottleneck, and it measures the bottleneck itself. Weisenmiller's report that Ali struggled particularly with the mathematical questions is consistent with this, since mid century arithmetic reasoning items are dense word problems.

This is a claim about test construct validity, not a claim about Ali's ability. The honest formulation is that a score obtained under those conditions is a valid measure of performance on that test on that day, and a poor measure of the broader construct the test is usually taken to represent. Modern practice recognises this, which is why accommodations exist and why the relationship between dyslexia and measured IQ is an active area rather than a settled one.

What this argument does not licenseReading difficulty explains why a written aptitude percentile is a narrow measure. It does not entitle anyone to substitute a higher number, including a flattering one. The correct conclusion is that the 1964 result does not support a general ability estimate in either direction, not that Ali's real score must have been higher. Replacing an unsupported low figure with an unsupported high figure is the same error twice.

One further caution belongs here. The specifics of any diagnosis, its date, and who made it are reported rather than documented in the sources reviewed for this page. Ali described his own difficulty publicly and that description is his to make. Retrospective clinical labelling of a person who cannot be assessed is not a diagnosis, and this page does not offer one.

10 What a four part aptitude composite leaves out

Even administered perfectly, to a fluent reader, with unlimited time and a modern reference group, the AFQT would still not be an IQ, because of what it never asks about.

The composite covers Arithmetic Reasoning, Mathematics Knowledge, Paragraph Comprehension and Word Knowledge. Mapped onto the framework used throughout contemporary ability research, that is a heavy load on Gq quantitative knowledge and Gc comprehension knowledge, and essentially nothing else. Whole domains are absent by design, because the armed services were screening for trainability in tasks that involve manuals and numbers, not building a cognitive profile.

CHC domainWhat it coversSampled by the AFQT composite
Gq quantitative knowledgeAcquired mathematical knowledge and its applicationYes, two of the four subtests
Gc comprehension knowledgeVocabulary, verbal knowledge, comprehension of textYes, two of the four subtests
Gf fluid reasoningNovel problem solving with unfamiliar materialOnly indirectly, through arithmetic reasoning items
Gv visual spatialMental representation and transformation of figuresNo
Gwm working memoryHolding and manipulating information across a short spanNo
Gs processing speedSpeed and fluency on simple, well learned operationsNo, though the whole paper is timed

A person can be strong in the two domains the AFQT samples and weak elsewhere, or the reverse, and the composite will not reveal it. That is precisely the information a profile is for. A modern battery reports a Full Scale figure alongside index scores so that a flat profile can be distinguished from a jagged one, because two people with identical Full Scale scores can have completely different patterns underneath, as the treatment of the six domains sets out.

ACIS is built on the other side of this trade. It administers 20 subtests across six primary cognitive domains, and reports Full Scale IQ alongside six index scores, with subtests on a mean 10 and standard deviation 3 scaled score metric and composites on the conventional mean 100 and standard deviation 15 metric. The measured breadth is why its Full Scale composite reaches an omega of .9886 and a g loading of .958 on a technical analysis set of 2,750 complete records, with a higher order confirmatory model fitting at CFI .9761, TLI .9726, RMSEA .0406 and SRMR .0217.

Those figures carry their own boundary and it is a real one. The sample is self selected rather than census based, the administration is unsupervised, and the adult reference frame is a modelled frame rather than a national probability sample. Strong internal structure is evidence that the battery is measuring something coherent and broad, which is the specific contrast with a four part screen. It is not evidence that an online test is a clinical instrument, and ACIS is not one. Readers deciding what an unsupervised result can carry should start at the comparison with supervised professional testing.

11 The overcorrection: eloquence is not a score either

The most common response to Ali's test result is to point at his verbal brilliance and conclude that the test was simply wrong about him. That instinct is generous, it is partly right, and as an argument it does not work.

Ali's verbal gift is not in dispute and does not need defending. He improvised rhyming verse at press conferences, controlled interviews with practised comic timing, and constructed public arguments on religion and the war that he sustained under hostile questioning for years. Whatever else the 1964 examination captured, it plainly did not capture that.

But a public performance is not an assessment, and treating it as one repeats the error the article has been arguing against, only with the sign flipped. Rhetorical fluency draws on verbal knowledge, working memory and social skill in combination, under conditions the performer chooses, on material they have rehearsed. A test administration is a standardised sample under fixed conditions. Neither substitutes for the other, and inferring a score from a biography is exactly the move that produces the unsupported numbers in the first place.

The defensible statement is narrower and stronger than the popular one. It is not that Ali's real IQ was high. It is that a timed written arithmetic and reading screen, referenced to a wartime population, administered to a man with documented reading difficulty, is too narrow an instrument to support any general ability claim about him at all. The result constrains almost nothing. Both the people who quote 78 to diminish him and the people who quote his poetry to refute it are claiming more than the evidence carries.

Ali's own account is the most graceful version of the point. Asked about the result, he said that he had claimed to be the greatest and not the smartest, and that when he looked at a lot of the questions he simply did not know the answers. That is a man declining to contest a finding about a narrow task while declining to accept it as a judgment on his worth, which is a more sophisticated position than most of the commentary written about him since.

He is also, for what it is worth, a poor example of the thing the genre wants him to illustrate. Group level findings about ability and life outcomes describe populations and predict individuals weakly, so a single famous case neither confirms nor refutes them. That asymmetry is why what a test measures and what a life demonstrates have to be argued separately.

12 The evidence ledger for every number on this page

Rather than leave the labels scattered, here is the complete accounting, so that a reader can see at a glance which claims survive scrutiny and which do not. The categories are the ones used throughout the site: documented, reported, attributed, estimated and unsupported.

Documented. That a mental qualifying examination was administered to Ali in 1964. That he was classified 1-Y in 1964 and reclassified 1-A in 1966 by Local Board No. 47. That the Supreme Court reversed his conviction for refusing induction in Clay v. United States in 1971. That the AFQT is composed of four subtests and reported as a percentile from 1 to 99. That the pre 1980 armed services interpreted recruit scores against World War Two era data, and that the battery in use from at least 1976 was misnormed in a way that inflated AFQT percentiles.

Reported. That his result fell at roughly the 16th to 18th percentile, given consistently by Page in 2016 and by Weisenmiller in 2015, who records a score of 16. That the qualifying floor was the 30th percentile in 1964 and the 15th from early 1966. That he had particular difficulty with the mathematical questions.

Attributed. The figure of 83, attributed by Page to Edmonds's 2006 biography and through it to FBI files, with the primary source unverified here. Ali's remarks about being the greatest rather than the smartest, attributed consistently across sources. His dyslexia and its role in the 1964 outcome, described by Ali himself and reported widely, without a clinical record in the public domain.

Estimated. That the 16th percentile corresponds to a standard score of about 85 on a mean 100 and standard deviation 15 scale. This is normal curve arithmetic and is valid only if both scales are referenced to the same population, which in this case they are not.

Unsupported. That Muhammad Ali's IQ was 78. That his IQ was 83. That his IQ was any single number. That the 1964 result establishes anything general about his cognitive ability. That his eloquence establishes a high score.

The one sentence versionMuhammad Ali sat a real, dated, official aptitude screen and placed below the cutoff then in force, and that fact supports no IQ point score whatsoever, because the instrument was narrow, the reference population was from another era, the conditions were adverse, and the output was a rank rather than a score.

Applying that discipline to a living person is the same exercise. If you want a number about yourself rather than about a public figure, the requirement is a full battery, a stated norm, a reported error term and an honest description of the conditions, which is what an accurate test is defined by and what adult testing should deliver. Researchers who need verified administrations for participants, with CSV export and no participant personal data, can run them through the ACIS Professional workspace, and the pitch for administering the battery to a group sits at the administration section on the home page.

13 How a testing professional would frame the whole case

Everything argued above is a restatement of published professional consensus, not a house position, and the consensus document is worth naming. The Standards for Educational and Psychological Testing, published in 2014 by the American Educational Research Association, the American Psychological Association and the National Council on Measurement in Education, is the reference work that governs this territory in the United States.

Its central move is to locate validity in a use rather than in a test. A test is not valid in the abstract. Evidence supports a particular interpretation of a particular score for a particular purpose, and changing the purpose requires new evidence. The AFQT carried validity evidence for screening applicants for military training. Nobody has ever assembled evidence supporting its use as a general intelligence estimate for one individual retold in a newspaper, and under the Standards that second use is unvalidated no matter how well the first was supported.

The Standards also require that a score report identify the norm group, that reliability information adequate for the intended interpretation be available, and that testing conditions and any factors likely to affect performance be considered when a score is interpreted. Ali's case fails all three. The norm group is a wartime service population never intended for civilian comparison, no reliability information accompanies the reported figures, and reading difficulty was present and unaccommodated.

They speak to fairness in the same breath. Fairness is treated as a validity issue rather than a courtesy: a score is not a fair measure of a construct if a characteristic irrelevant to that construct, such as impaired access to printed text, materially depresses performance. That is a technical judgment about what the instrument measured, and it holds regardless of anyone's view of the man.

Applying the Standards honestly means applying them to ACIS as well. ACIS is an unsupervised online assessment for adults aged 16 to 90. It is not a clinical or diagnostic instrument, and it is not appropriate for diagnosis, hiring decisions, educational accommodations or high IQ society admission. Its normative frame is self selected rather than census based. Its reliability and structural figures are strong and are published with those boundaries attached rather than despite them, which is the whole content of the published technical documentation and of the discussion of how measurement quality is established.

The reason to hold that line on a boxer who died in 2016 is that the same reasoning protects everyone who takes a test today. A number without a named instrument, a stated norm group and an error band is not a measurement of a person. It is a rumour with a decimal point. Muhammad Ali sat a real examination, and the most respectful and the most accurate thing that can be said about it is that it tells us what he scored on a narrow screen on one day in 1964, and nothing at all about the size of his mind.

14 Frequently Asked Questions

What was Muhammad Ali's IQ?

There is no documented IQ score for Muhammad Ali, because no IQ test publisher has reported administering a battery to him. What exists is a 1964 armed forces aptitude screen on which he was reported to place at roughly the 16th percentile.

Was Muhammad Ali's IQ 78?

The figure of 78 is widely repeated in press accounts but arrives without a named instrument, date or reference group. It also cannot be his AFQT percentile, since the reported percentile was 16 and the AFQT is scored on a 1 to 99 percentile scale.

Where does the figure of 83 come from?

Columnist Clarence Page reported it on 8 June 2016, attributing it to Anthony O. Edmonds's 2006 biography of Ali and to material in FBI files. The book is real and catalogued, but the underlying record has not been inspected for this page.

Which of the two numbers is right, 78 or 83?

Possibly neither. Two different point scores circulating for one sitting means at least one is wrong, and the likeliest explanation is that both are later reconstructions of a record that only ever contained a percentile.

What test did Muhammad Ali actually take?

He sat the written mental qualifying examination used by the United States armed forces for pre induction screening. It is an eligibility and trainability screen, not a clinical intelligence battery.

What does the AFQT measure?

Its modern form is a composite of four subtests: Arithmetic Reasoning, Mathematics Knowledge, Paragraph Comprehension and Word Knowledge. Research programmes describe it as a general measure of trainability and a criterion of enlistment eligibility.

Is the AFQT an IQ test?

No. It is scored as a percentile rank against a reference group rather than as a standard score, it was designed to predict trainability for military work, and it samples only quantitative and verbal content.

What is the difference between a percentile and an IQ score?

A percentile is an ordinal rank saying how many people you finished at or above. An IQ is a standard score saying how far you sat from a mean of 100 in fixed units of 15. The percentile chart on this site lists the correspondence.

Does the 16th percentile convert to an IQ of about 85?

Only as arithmetic, and only if both scales are referenced to the same normally distributed population. Ali's percentile was referenced to a mid century service population, so the conversion does not hold for him.

Why was Ali classified 1-Y in 1964?

Because his score on the mental qualifying examination fell below the qualifying floor then in force, reported as the 30th percentile. Class 1-Y meant acceptable for service only in a declared national emergency.

How did he become draft eligible in 1966 without retaking the test?

The Department of Defense lowered the mental qualification floor to the 15th percentile in early 1966. His unchanged result then sat above the new line and his local board reclassified him 1-A.

Does that mean the test was wrong?

No, it means the threshold was a policy choice rather than a property of the test or of Ali. A perfectly reliable score can sit either side of a line that an institution moves for reasons of manpower.

Did Muhammad Ali have dyslexia?

He described lifelong reading and writing difficulty publicly and named it himself in his later memoir, and reference sources describe him as dyslexic. The specifics of any formal diagnosis are reported rather than documented in the public record.

How would reading difficulty affect a written aptitude test?

A timed paper test filters every question through decoding. For someone whose access to printed text is impaired, the result measures performance on that task under those conditions rather than the broader construct the test is taken to represent.

Does his reading difficulty mean his real score was higher?

That does not follow. Adverse conditions mean the result cannot support a general ability claim in either direction. Substituting a flattering number for an unflattering one repeats the same error with the sign reversed.

Why does the 1964 reference population matter so much?

Because a percentile is meaningless without naming the group it ranks against. Before 1980 the services interpreted recruit scores against World War Two era data, so a 1964 percentile is not a rank among contemporaries.

Have military test norms ever been shown to be wrong?

Yes. A National Academies workshop report records that the battery in use from at least 1976 was misnormed in a way that inflated AFQT percentiles, and that many applicants who did not meet the standard were enlisted before the error was caught.

Doesn't Ali's eloquence prove he was highly intelligent?

His verbal gift is not in dispute, but a public performance under chosen conditions is not a standardised assessment. Inferring a score from a biography is the same move that produced the unsupported numbers this page is correcting.

What did Ali himself say about the result?

He said he had claimed to be the greatest rather than the smartest, and that when he looked at many of the questions he simply did not know the answers. He later campaigned publicly on literacy.

Has ACIS ever tested Muhammad Ali?

No. ACIS has never administered any assessment to Muhammad Ali, and no figure on this page is an ACIS measurement or an ACIS estimate of his ability. Every number is attributed to the source that reported it.

What would a defensible IQ estimate actually require?

A full battery with breadth across cognitive domains, a named and current norm group, a reported standard error so the score arrives as a band, and testing conditions documented well enough to judge what the score measured.

Take the assessment

You get a profile, not a number

ACIS measures six CHC domains across 20 subtests and reports each one with its own normed score and confidence interval, so you can see where you are strong and where you are not.

Free trial, no card required. Full report from $15.