Testing History

Adolf Hitler's IQ: the score that was never taken

ACIS has never assessed Adolf Hitler, and no retrospective assessment of any person is possible with any instrument. He died on 30 April 1945, months before the psychological testing program at the Nuremberg jail began, so no administration exists and none ever did. This page is about the testing that did happen: what was administered to the defendants, by whom, under what conditions, and what the examiners themselves recorded about the limits of their own numbers.

A black and white photograph of Adolf Hitler in a belted trench coat with his right arm extended, buildings behind him.
He died on 30 April 1945, months before the psychological testing program at the Nuremberg jail began.

0 Quick Answer

There is no IQ score for Adolf Hitler, because no intelligence test was ever administered to him, and the chronology makes it impossible that one could have been. He died on 30 April 1945. The International Military Tribunal at Nuremberg opened on 20 November 1945, and the psychological testing of the defendants was completed before that date. Every figure circulating for him, whether 135, 140, 141 or 145, was produced by someone estimating rather than measuring. ACIS has not assessed him, and no test can be given to a person retrospectively.

What does exist is one of the most documented episodes in the history of psychological assessment. Douglas Kelley, the psychiatrist assigned to the Nuremberg jail, and Gustave Gilbert, the prison psychologist and interpreter, administered an intelligence test and projective personality tests to the men who did stand trial. Gilbert published the intelligence results as a table in Nuremberg Diary in 1947, with 21 figures running from 143 for Hjalmar Schacht down to 106 for Julius Streicher. Kelley published his own account, 22 Cells in Nuremberg, the same year.

The reason this page is a history of testing rather than a curiosity is that those numbers cannot be read the way they are usually read. Gilbert built a German adaptation of an American instrument, removed the subtests most affected by language and culture, applied an age allowance, and then printed a warning that four of the older men would score 15 to 20 points lower on an absolute basis than his table shows. He also wrote the sentence that matters most on any page of this kind: that the IQ "indicates nothing but the mechanical efficiency of the mind, and has nothing to do with character or morals."

30 April 1945

Date of Hitler's death, before any Nuremberg testing was carried out.

21 figures

Intelligence test results Gilbert printed in Nuremberg Diary, one for each defendant who stood trial in person.

15 to 20 points

The downward adjustment Gilbert himself said applied on an absolute basis to the four oldest men in his table.

What this page is and is notThis is an account of an episode in the history of psychological measurement. It does not offer an estimate of anyone's ability, does not present any score as a credit or a discredit to the person who produced it, and does not treat a test result as an explanation of conduct. Measured cognitive ability carries no moral content, and the Nuremberg testing is studied precisely because it addressed that question directly.

1 The Chronology That Closes the Question

The absence of a score here is not an archival gap that might one day be filled. It is a consequence of dates. Setting them out in order is the fastest way to see why no future discovery will produce a Hitler test record.

He was born on 20 April 1889 in Braunau am Inn, according to the Holocaust Encyclopedia maintained by the United States Holocaust Memorial Museum, and died in 1945. The date of death is 30 April 1945. The International Military Tribunal opened at Nuremberg on 20 November 1945. Gilbert states in Nuremberg Diary that "the psychological testing program was virtually completed before the beginning of the trial while the prisoners were still in solitary confinement so that the validity of the results was safeguarded." The testing therefore ran in the autumn of 1945, roughly six months after his death.

The chronology matters for a second reason. There was no earlier opportunity either. Intelligence testing as a routine practice on adults barely existed in German speaking Europe during the years he could have been tested. The Stanford revision of the Binet scale was published in 1916 and was built for children. The instrument used at Nuremberg was itself published only in 1939. The window in which an adult in Central Europe would plausibly have sat a standardized intelligence test, outside military or clinical selection, was very narrow and he was never in it.

The third point is about the population of records that does exist. School reports, army records, admission decisions and medical notes survive for many figures of that period, and historians work with them. None of those documents is an intelligence test. A test result is a specific artifact: a named instrument, a date, an examiner, a raw total, and a conversion against a reference sample. The absence of that artifact is not a matter of interpretation, and no biography can substitute for it. The broader pattern, where circulating figures for famous names dissolve on inspection into estimates and repetitions, is documented on the page on record high IQ claims.

2 What Was Actually Administered at Nuremberg

Two men ran the psychological program at the Nuremberg jail, they had unusual access, and they left contradictory published accounts of what it showed. Getting their roles right is necessary before any of the numbers make sense.

Douglas M. Kelley was the psychiatrist assigned to the jail. Gustave M. Gilbert was the prison psychologist. In 22 Cells in Nuremberg, Kelley describes the division of labor plainly: "The intelligence estimations were made from an adaptation of the Wechsler-Bellevue Test devised by my assistant, Dr. Gustave Gilbert, the prison psychologist and a captain in the A.U.S." He adds that Gilbert was also assigned to his office as an interpreter and, at Kelley's direction, made records of many of his conversations with the prisoners.

The battery had three parts. Gilbert lists them in order of how simply they can be described: an intelligence test, the Rorschach inkblot test, and the Thematic Apperception Test. He notes in the same passage that the intelligence test, though the easiest to summarize, was not the most significant part of the program, which is a judgment worth keeping in mind given that it is the only part anyone quotes.

The conditions of administration were unusual in both directions. On one side, the examiners had access no clinician normally gets: the men were confined, available daily, and had every incentive to cooperate with the medical staff. On the other, they were defendants in a capital proceeding, held in solitary confinement, several of them elderly, all of them under conditions Kelley describes as producing depression of varying degrees in all of the prisoners. Conditions of administration are part of what a score means, and neither examiner pretended otherwise.

ElementWhat the primary sources record
ExaminersDouglas M. Kelley, psychiatrist to the jail; Gustave M. Gilbert, prison psychologist and interpreter, a captain in the A.U.S.
Intelligence instrumentGilbert's own German version of the American Wechsler-Bellevue Adult Intelligence Scale, adapted
Other instrumentsRorschach inkblot test and Thematic Apperception Test
TimingVirtually completed before the trial opened on 20 November 1945
ConditionsPrisoners in solitary confinement, awaiting a capital trial
Where the results were publishedGilbert, Nuremberg Diary, 1947; Kelley, 22 Cells in Nuremberg, 1947

One defendant is absent from the intelligence table. Robert Ley died by suicide in his cell in October 1945, before the trial opened, and Kelley describes finding him. Gilbert's table therefore carries 21 names rather than the 22 prisoners Kelley writes about.

3 The Scores Gilbert Recorded and Published

The table below is transcribed from page 34 of Nuremberg Diary and is the only published intelligence record from the Nuremberg proceedings. It is reproduced here as a historical document, in the order and with the figures Gilbert printed. The numbers are not endorsements, comparisons, or evaluations of anyone. They are what one examiner computed with one adapted instrument in one autumn.

DefendantFigure printed by GilbertDefendantFigure printed by Gilbert
Hjalmar Schacht143Albert Speer128
Arthur Seyss-Inquart141Alfred Jodl127
Hermann Goering138Alfred Rosenberg127
Karl Doenitz138Constantin von Neurath125
Franz von Papen134Walther Funk124
Erich Raeder134Wilhelm Frick124
Hans Frank130Rudolf Hess, estimated120
Hans Fritzsche130Fritz Sauckel118
Baldur von Schirach130Ernst Kaltenbrunner113
Joachim von Ribbentrop129Julius Streicher106
Wilhelm Keitel129

Three details in that table are usually dropped when it is reproduced. The Hess figure carries an asterisk in the original, and Gilbert's footnote reads that it is given on the basis of a retest after recovery. Kelley, writing about the same man in the same year, gives a different account: he reports that the Rorschach disclosed an IQ of between 115 and 120 and that this was confirmed by a straight intelligence test given by the prison psychologist. Two books published in 1947 by the two men who ran the program report the same person as 120 and as a range of 115 to 120.

The second detail is that Gilbert put a floor under his own interpretation. His comment on the table reads that except for Streicher, the figures show the defendants above the average band he defines as 90 to 110, and he immediately characterizes this as merely confirming that the most successful men in any sphere of activity, whether politics, industry, militarism or crime, tend to be above average. That is a statement about selection into positions of power, and it is the only inference he draws from the ranking.

The third is that the figures are not the part of the program Gilbert considered important. He says so in the paragraph introducing them. The Rorschach and Thematic Apperception records, which are the subject of the scholarly literature covered further down this page, were what the psychological program existed to produce.

4 The Caveats the Examiner Printed Beside His Own Table

Gilbert published his methodological limits on the same page as his results, and every one of them narrows what the numbers can support. Reading the table without them is the single most common error in the material that circulates.

The first caveat is about the instrument. In his own words: "I used my own German version of the American Wechsler-Bellevue Adult Intelligence Test, eliminating and compensating for those parts which are subject to cultural differences, like vocabulary and general information." He then lists the eight tasks that remained. The verbal group was memory span for number series of increasing length, simple arithmetic of increasing difficulty, common sense questions, and concept formation by verbal similarities. The performance group was a substitution code test, object assembly, converting designs on colored blocks, and recognizing missing parts of pictures.

The second caveat is about age. Gilbert notes that the figures were calculated by the Wechsler-Bellevue method, which unlike the Stanford-Binet makes allowance for decline in average performance with age rather than assuming a constant adult level, and that this gave a fairer comparison across men of widely divergent ages. Then comes the sentence almost nobody reproduces: "It should be borne in mind, however, that the effective intelligence of older men like von Papen, Raeder, Schacht, and Streicher was 15-20 points lower than the IQ's indicated here but their relative standing in their respective age groups is accurately indicated."

Read that against the table. The top figure of 143 belongs to one of the four men Gilbert flagged. So does the bottom figure of 106. The rank order that circulates as a curiosity is, by the examiner's own statement, a set of within age standings and not a common absolute scale, and the two most quoted entries in it are both inside the group he singled out.

8 tasks

Subtests remaining in Gilbert's adapted battery after he removed vocabulary and general information.

4 of 21

Men Gilbert named as carrying an absolute adjustment of 15 to 20 points, including both the highest and the lowest figures.

90 to 110

The average band Gilbert used when interpreting his own table.

The third caveat is the one this whole page turns on, and it is Gilbert's, not ours. He writes that the IQ "indicates nothing but the mechanical efficiency of the mind, and has nothing to do with character or morals, nor the various other considerations that go into an evaluation of personality." A psychologist who had spent months alone with these men, and who would spend the rest of his career arguing that they shared a describable pathology, put that boundary in print on the page that carries his scores.

5 What the Wechsler-Bellevue Was, and Whose Distribution It Described

The instrument was six years old at the time, it was built for an American reference population, and it was the first widely used scale to abandon mental age arithmetic. Each of those three facts changes what the Nuremberg figures can be compared to.

The Wechsler-Bellevue Intelligence Scale, Form I, was published in 1939 and named for David Wechsler and the Bellevue Hospital in New York where he was chief psychologist. The Science Museum Group catalogue record for a surviving set dates the materials to 1939 onward, published in New York by The Psychological Corporation, and lists the physical components: picture completion booklets, block design cards, arithmetic problems, object assembly puzzles, picture arrangement cards and digit symbol materials. Those are recognizably the tasks Gilbert kept.

Its methodological importance is the scoring change. Earlier scales computed a ratio: mental age divided by chronological age, multiplied by 100. That formula was built for children and misbehaves for adults, because mental age stops being a coherent unit once growth flattens. Wechsler's scales rank a person against a reference sample of adults and express the result as a distance from that group's average. Gilbert's remark that his method, unlike the Stanford-Binet, allows for decline with age rather than assuming a constant adult level is that distinction stated from inside the period. The transition from one system to the other is set out on the mental age page, and the arithmetic of both is on the page on how IQ is calculated.

The consequence for Nuremberg is direct. A deviation score is meaningful only relative to the sample it was normed against, and that sample was American and pre-war. When the same instrument is translated, adapted, and administered to German defendants in 1945, the resulting number still refers to an American reference distribution from the 1930s. It answers the question of where these men would fall against that group if everything else transferred cleanly. Everything else did not transfer cleanly, which is the next two sections.

None of this means the figures were badly produced. Gilbert did roughly what a careful examiner in 1945 could do with the tools available, and he documented what he did, which is more than most published score tables of any era offer. The point is about what a number can carry, not about the competence of the person who computed it. The general dependence of any score on its normative frame is covered on the page on norming.

6 What Translation and Adaptation Do to a Score

A translated and abridged test is a different instrument from the one it was translated from, and its scores are not interchangeable with the original's. This is settled in the modern test adaptation literature and it was already visible in what Gilbert wrote about his own procedure.

Start with what he removed. Vocabulary and general information are the two Wechsler subtests most saturated with a specific language and a specific culture's stock of facts, and they are also among the most reliable indicators of crystallized verbal ability in the whole battery. Dropping them is defensible when you are testing across a language barrier. It also means the remaining battery samples a narrower slice of ability, weighted toward working memory, arithmetic, and visual and constructional tasks, and away from accumulated verbal knowledge.

On the ACIS published figures, the difference in what those task families measure is quantified. Verbal Comprehension has a g loading of .864 and an omega reliability of .9745 across five indicators, while Working Memory has a g loading of .788 and Processing Speed .648. The two are not redundant either: Verbal Comprehension correlates .577 with Processing Speed, one of the lowest index intercorrelations in the battery. A composite that drops the verbal knowledge indicators and keeps the memory and constructional ones is not the same composite with a bit less precision. It is weighted differently, and it estimates a different mix. Those figures come from the normative model published in the technical manual.

Then there is the language of administration itself. Kelley records that he always used interpreters to avoid misunderstanding in important discussions, and that he would rotate interpreters on different days to get the same information through more than one channel. That is careful practice for clinical interviewing. For standardized testing it introduces a variable the norms do not contain, because the reference sample was tested in English by examiners working from a printed script. Every departure from the standardized administration widens the gap between the score and the frame it is being read against.

What cannot be recoveredNo correction factor exists that would convert Gilbert's figures onto a modern scale, or onto any contemporary German reference frame, and none can be constructed now. The item level records are not available, the adaptation was his own and unpublished as an instrument, and there is no German standardization sample from 1945 to convert against. Anyone presenting these numbers as comparable to a present day score is asserting something the evidence cannot support.

The problem Gilbert was solving is still a live problem. Building an ability test that does not smuggle in one culture's vocabulary is the explicit design goal of a whole family of instruments, and the tradeoffs involved are set out on the page on culture fair testing.

7 Why a 1930s Reference Frame Does Not Transfer to Now

Even if the translation problem vanished, a figure computed against a 1930s American norm would still not be readable next to a modern score, because population test performance has moved. This is the best measured phenomenon in the whole area and it has a formal meta-analysis behind it.

Pietschnig and Voracek reported in Perspectives on Psychological Science, volume 10, issue 3, pages 282 to 306, in 2015, the first formal meta-analysis of the Flynn effect. Across 271 independent samples from 31 countries, totaling almost 4 million participants and spanning 1909 to 2013, they estimated annual gains of 0.41 IQ points for fluid test performance, 0.30 for spatial, 0.28 for full scale, and 0.21 for crystallized. They also found the gains stronger for adults than for children and decreasing in more recent decades.

Apply the full scale figure mechanically to the interval between a 1939 standardization and today and you get a very large number, which is exactly why the mechanical application is wrong. A meta-analytic trend describes average movement across many samples and instruments. It is not a conversion table for one administration, it says nothing about which direction a specific 1945 German cohort would sit relative to a 1930s American frame, and the domain differences in the same paper show that the size of any adjustment would depend on which subtests were in the battery. Since Gilbert removed the two most crystallized subtests, even the domain weighting is unknown.

The honest statement is that the direction of the bias is unknown and its magnitude is unknowable. What is known is that norm frames age, that publishers restandardize their instruments for exactly this reason, and that comparing a score across eight decades of norm generations is not a small approximation. The mechanism and its consequences are set out on the Flynn effect page, and what a standard deviation of 15 does and does not fix is on the page on the 15 point scale.

Obstacle to comparisonWhat it does to the Nuremberg figuresCan it be corrected now
American reference sample, pre-warThe score describes standing against a group that is not the test taker's populationNo, no contemporary German frame exists to convert against
Translation into German by the examinerIntroduces variance the standardization did not containNo, the adaptation was never published as an instrument
Vocabulary and information subtests removedShifts the composite away from crystallized verbal abilityNo, item level records are unavailable
Age allowance applied by the examinerMakes figures within age standings rather than a common absolute scalePartly, Gilbert stated the size for four men
Eight decades of norm driftPlaces the figures on a frame no current score sharesNo, meta-analytic trends are not conversion factors
Administration under capital detentionConditions differ from the standardization protocolNo, and the direction of the effect is not established

8 The Disagreement Between the Two Examiners

Kelley and Gilbert ran the same program on the same men and published opposite conclusions about what it showed, and that dispute, not the intelligence table, is the reason the Nuremberg records still appear in the psychological literature. Both positions are in print and both should be stated.

Kelley's conclusion in 22 Cells in Nuremberg is that the defendants were not a distinct psychological type. He opens the book by rejecting the framing he had been met with on returning to the United States: "Insanity is no explanation for the Nazis. They were simply creatures of their environment, as all humans are." He closes it by generalizing the point. His summary reads that the Nazi leaders "were not spectacular types, not personalities such as appear only once in a century," and that they "simply had three quite unremarkable characteristics in common, and the opportunity to seize power," which he lists as overweening ambition, low ethical standards, and a nationalism that justified anything done in its name. He argues in the same passage that personalities of the kind he examined "can be found anywhere in the country, behind big desks deciding big affairs as businessmen, politicians, and racketeers."

Gilbert took the opposite view and spent his subsequent career on it. Where Kelley saw ordinary men in extraordinary circumstances, Gilbert argued that the group displayed a describable and shared pathology, and the two broke over the interpretation. Their disagreement extended to individual cases, including Streicher, whom Gilbert regarded as showing a paranoid pattern and Kelley regarded as essentially rational and fixated on a single obsession. Gilbert introduced and supplied the Rorschach records for the 1975 volume that argued hardest for the pathology reading.

Two things are worth noticing about the structure of that argument. The first is that the disagreement was not about the data, which both men had, but about what the data licensed, which is the recurring problem in projective testing. The second is that Kelley's position is the more uncomfortable one and it is the one the later evidence has tended to support, which is the subject of the next section.

Both men also went further than their instruments allowed at times, in opposite directions. That is not a modern complaint imposed on them. It is what the researchers who reanalyzed their records concluded, working from the same protocols with better scoring systems and thirty to fifty years of distance.

9 What Fifty Years of Reanalysis Found in the Records

The Nuremberg Rorschach protocols have been reanalyzed at least five times by different research groups, and the weight of those analyses runs against the idea of a single Nazi personality. The literature is small, traceable and largely in one journal, which makes it unusually easy to follow.

Florence Miale and Michael Selzer published The Nuremberg Mind in 1975, with an introduction and the Rorschach records supplied by Gilbert. It presents the responses of 16 defendants with interpretive analyses and argues for a distinctive pathology. It is the strongest published statement of the position Gilbert held.

Molly Harrower tested that position experimentally the following year. Reporting in the Journal of Personality Assessment, volume 40, issue 4, pages 341 to 351, in 1976, she took the records of 17 Nazi war criminals administered in 1946 by Kelley and Gilbert, selected eight of them, matched them with eight control records for level of mental health potential, and had ten Rorschach experts assess all sixteen blind. Her reported result is that the Nazi records were not identified as such. The experts sorted the protocols by adjustment and inadequacy across both groups, which is what you would expect if the two sets were not distinguishable on the dimensions the raters were using.

Barry Ritzler followed with a quantitative approach in the same journal, volume 42, issue 4, pages 344 to 353, in 1978. Eric Zillmer, Robert Archer and Ruth Castino then reanalyzed the records using Exner's Comprehensive Scoring System and reported in the Journal of Personality Assessment, volume 53, issue 1, pages 85 to 99, in 1989, that earlier studies had made two kinds of error: overinterpretation and excessive inference, and a failure to detect meaningful distinctions between protocols representing different personality styles. Their conclusion is stated directly in the abstract, that the records "cannot be grouped together into one specific mental disorder that would adequately characterize these diverse individuals."

Two later papers extended the argument in both directions. Resnick and Nunno attempted a blind actuarial analysis using the Comprehensive System in the Journal of Personality Assessment, volume 57, issue 1, pages 19 to 29, in 1991. Greiner and Nunno compared 16 of the records against those of incarcerated men diagnosed with antisocial personality disorder in the Journal of Clinical Psychology, volume 50, issue 3, pages 415 to 429, in 1994, and reported that the Nuremberg records matched the standard psychopathy hypotheses less well than the comparison inmates did, and that variance in type and degree of psychopathology precluded applying any single characterization to most of them.

Joel Dimsdale's review in the Journal of Psychosomatic Research, volume 78, issue 6, pages 515 to 518, in 2015, records where the question stands. He notes that the observations were kept from public view for decades and that there remains controversy even now about what the records revealed. That is an accurate summary. The longer treatment is Zillmer, Harrower, Ritzler and Archer, The Quest for the Nazi Personality, published by Lawrence Erlbaum in 1995.

10 A Diagnosis With No Examination Behind It

The clearest illustration of the method problem on this page comes from one of the examiners themselves, who published a psychiatric characterization of a man he had never met. It is worth setting out because it shows the failure mode operating at the highest level of access anyone has ever had to this material.

Kelley devotes a section of 22 Cells in Nuremberg to Hitler, and it contains a formal sounding conclusion: "In simple terms, Hitler was an abnormal and a mentally ill individual, though his deviations were not of a nature which in the average individual would arouse the serious concern of others." The preceding paragraphs assemble that judgment from reported symptoms, described behavior, and the accounts of associates.

The inputs are exactly what a modern reader should notice. Kelley never examined him, administered nothing to him, and had no test protocol, no interview and no clinical observation. He had documents, interviews with the man's surviving colleagues, and the professional confidence of a psychiatrist in 1947. He was working from a biography, which is the same evidence base every circulating IQ figure for the same person rests on, and he was working with far better access to that biography than any later writer.

The point is not that Kelley was careless. He was a competent clinician writing at the outer edge of what his sources allowed, and he says elsewhere in the book that outside evidence let him verify virtually every facet of character. The point is that access to a biography, however good, does not produce a measurement, and the professional standing of the person doing the inferring does not change that. Once you accept a retrospective characterization from a document set, the difference between a psychiatrist in 1947 and a website in 2026 is one of quality, not of kind.

That is the historiographic hinge for this whole subject. The Nuremberg intelligence figures exist because someone sat in a room with these men and administered a test. Everything said about the one man who was not in that room, by Kelley then and by aggregator pages now, belongs to a different evidence class entirely. Keeping the two apart is the only thing that makes the record usable.

The boundary this page holdsA published characterization by a named clinician who examined nobody is a documented opinion, not a finding. It is reported here because it is part of the historical record of the Nuremberg psychological program and because it demonstrates the method under discussion, not because it establishes anything about the person it describes.

11 Where the Circulating Figures Come From

The numbers attached to Hitler online, most often 135, 140, 141 and 145, do not trace to any administration, any examiner, or any document. They trace to forum threads, ranked list pages and estimate sites, each of which cites the previous one.

Follow any of them backward and the trail ends the same way. There is no test date because there was no test. There is no examiner because nobody examined him. There is no instrument named, no raw score, no reference sample, and no document to open. What there is instead is a figure that appeared somewhere, got copied, and became stable precisely because nothing anchors it. A claim with no source cannot be corrected by checking the source.

The underlying method, where one exists at all, is estimation from biography, and its properties are known because its inventor measured them. Catharine Cox estimated childhood IQ for 301 eminent historical figures in the 1926 volume of Terman's series, working from case histories assembled out of more than 1,500 biographical sources. She discovered after tabulating everything that the more complete the surviving record, the higher the estimate her raters produced, a relationship she described as entirely unexpected. Her estimates correlated .25 with rank order of eminence, and .16 once the reliability of the underlying data was held constant, inside a group already selected for fame. In other words a large part of what the method measures is how much paper survived. The full account of Cox's volume, its correction procedure and her own warnings about it is on the page on the genius threshold.

Estimates about the living and recently dead are worse still, because the surviving documentation was produced by people with something at stake in how the subject is remembered. The general weakness of estimates, including the finding that self estimated intelligence correlates only about .30 with measured ability across 93 studies and 36,833 participants, is set out on the page on estimating your own IQ. The same evidence class problem, worked through for a living political figure with a large documentary record and no test result, is on the Joe Biden page.

The documentary record that does exist for Hitler is a record of institutional decisions. The Holocaust Encyclopedia records that his family moved to Linz in 1898, that he sought a career in the visual arts against his father's wishes, and that he lived in Vienna from February 1908 to May 1913 supporting himself by painting watercolors and sketches. Historians working from surviving admission records report that he applied twice to the Vienna Academy of Fine Arts and was refused both times. A school report, an admissions decision and a rejection letter are documents about institutions and about what those institutions chose to do. None of them is a score, none of them was produced by a standardized instrument, and none can be converted into one.

12 What a Score Would Have Required, Then and Now

Four things are needed before a number describing a person's measured ability means anything, and this case is missing all four. Setting them out is more useful than another paragraph on what is absent.

  • A named instrument with a published manual. Not a description of a test, but a specific edition with documented content and documented norms. The Nuremberg record has this, in adapted form. The Hitler record has nothing.
  • A date, an examiner and recorded conditions. Gilbert supplied all three, including the fact that the men were in solitary confinement. This is the part most historical score claims omit entirely.
  • A reference sample the raw performance was converted against. Gilbert's was American and from the 1930s, which is why his figures are readable within his table and not outside it.
  • A document a third party can open. Nuremberg Diary and 22 Cells in Nuremberg are both scanned and public. A figure that lives only in the citation chain of other pages fails this test by construction.

Measured against that list, the Nuremberg intelligence table is a real historical record with severe and self documented limits, and the figures circulating for Hitler are not records at all. Those are different failures and they should not be described in the same language. The distinction between a documented result, a reported one and an attributed one runs through every page in this section, including the audit sorting celebrity figures by evidence class and the catalogue of claims that survived by repetition.

One further boundary belongs here, stated flatly. A cognitive score describes performance on a set of standardized tasks relative to a reference group. It carries no moral content, in either direction, and it never has. Gilbert wrote that on the page carrying his own table in 1947, and the reason the Nuremberg records are still studied eighty years later is that the program was designed to test whether measurable psychological characteristics distinguished the men in those cells from everyone else. The answer the reanalysis literature has converged on is that they did not, and that finding is the durable result of the whole episode.

For a reader whose interest is their own profile rather than a historical one, the relevant instruments are the modern normed batteries, and what they measure is set out on the page on what IQ tests measure and the inventory of instruments in current use. ACIS is an unsupervised online assessment for adults aged 16 to 90, running 20 subtests across six CHC domains. Its Quick form covers six subtests in about 45 minutes for 15 dollars and reports Verbal Comprehension, Fluid Reasoning and Working Memory. It is not a clinical instrument and is not appropriate for diagnosis, hiring decisions or accommodation requests. The subtests it uses are described on the Quick form page.

13 Sources Behind This Page

Every figure and every quotation above comes from a document that can be opened, and the two primary books are scanned and public. The list is given in full so the page can be checked against its own sources.

  • The primary intelligence record. Gustave M. Gilbert, Nuremberg Diary, 1947. The table of 21 figures, the description of the adapted battery, the age caveat and the statement that the IQ has nothing to do with character or morals are all on and around page 34 of the scanned edition.
  • The psychiatrist's account. Douglas M. Kelley, 22 Cells in Nuremberg, 1947. Source of the description of Gilbert's role and adapted instrument, the use and rotation of interpreters, the Hess figure of 115 to 120, the characterization of Hitler, and the concluding argument that the defendants were not a distinct type.
  • The pathology reading. Florence R. Miale and Michael Selzer, The Nuremberg Mind: The Psychology of the Nazi Leaders, 1975, with an introduction and Rorschach records by Gustave M. Gilbert.
  • The blind experiment. Molly Harrower, Journal of Personality Assessment, 40(4), 341 to 351, 1976, doi 10.1207/s15327752jpa4004_1.
  • The quantitative reanalysis. Barry A. Ritzler, Journal of Personality Assessment, 42(4), 344 to 353, 1978, doi 10.1207/s15327752jpa4204_2.
  • The Comprehensive System rescoring. Eric A. Zillmer, Robert P. Archer and Ruth Castino, Journal of Personality Assessment, 53(1), 85 to 99, 1989, doi 10.1207/s15327752jpa5301_10.
  • The actuarial attempt. M. N. Resnick and V. J. Nunno, Journal of Personality Assessment, 57(1), 19 to 29, 1991, doi 10.1207/s15327752jpa5701_3.
  • The psychopathy comparison. N. Greiner and V. J. Nunno, Journal of Clinical Psychology, 50(3), 415 to 429, 1994.
  • The historical review. Joel E. Dimsdale, Journal of Psychosomatic Research, 78(6), 515 to 518, 2015, doi 10.1016/j.jpsychores.2015.04.001.
  • The Flynn effect meta-analysis. Jakob Pietschnig and Martin Voracek, Perspectives on Psychological Science, 10(3), 282 to 306, 2015, doi 10.1177/1745691615577701.
  • The biographical estimation method. Catharine Cox, The Early Mental Traits of Three Hundred Geniuses, Stanford University Press, 1926.
  • Historical background. United States Holocaust Memorial Museum, Holocaust Encyclopedia.
  • The instrument. Science Museum Group collection record for Wechsler-Bellevue Intelligence Scale Form 1 test material, published in New York by The Psychological Corporation from 1939.

Three things on this page could not be independently confirmed and are marked as such in the text. The exact standardization sample size of the Wechsler-Bellevue Form I is not stated here because no primary source for it was available. The precise dates on which each defendant was tested are not recorded in either published account beyond Gilbert's statement that the program was virtually complete before the trial opened. The Vienna Academy applications are attributed to historians working from surviving admission records rather than to a document reproduced here.

The professional framework is explicit and it applies to historical records as much as to current ones. The Standards for Educational and Psychological Testing (2014), published jointly by the American Educational Research Association, the American Psychological Association and the National Council on Measurement in Education, require that a score interpretation be supported by validity evidence for the specific use proposed, that the characteristics and limits of the normative sample be disclosed, that any adaptation or translation of an instrument be documented and separately validated, and that scores be reported with their measurement error. APA standards on test use add that an instrument may not be used for a purpose its validity evidence does not cover. A translated and abridged battery normed on a different population eight decades ago, read as though it produced scores comparable to a modern administration, fails several of those requirements at once. Gilbert, writing before any of those standards existed, printed most of the relevant warnings himself.

14 Frequently Asked Questions

What was Adolf Hitler's IQ?

No intelligence test was ever administered to him, so no score exists. He died on 30 April 1945, months before the psychological testing program at the Nuremberg jail began, and there is no record of any earlier testing.

Where do the figures of 135, 140 and 141 come from?

From forum posts, ranked list pages and estimate sites that cite one another. None names an instrument, a date, an examiner or a document, which is what distinguishes an attributed figure from a documented one.

Has ACIS assessed him?

No. ACIS has never assessed him and no retrospective assessment of any person is possible with any instrument. Nothing on this page is an ACIS measurement or an ACIS estimate.

Who was tested at Nuremberg?

The defendants who stood trial in person. Gustave Gilbert, the prison psychologist, printed 21 figures in Nuremberg Diary in 1947. Robert Ley, who died by suicide in his cell in October 1945, does not appear in the table.

What test was used?

Gilbert's own German adaptation of the Wechsler-Bellevue Adult Intelligence Scale, an American instrument published in 1939. He states that he eliminated and compensated for the parts most subject to cultural differences, naming vocabulary and general information.

What was the highest score recorded?

Hjalmar Schacht at 143, the top entry in Gilbert's table. Schacht is also one of the four older men Gilbert singled out as carrying an absolute adjustment of 15 to 20 points downward.

What was the lowest?

Julius Streicher at 106, the bottom entry. He is also inside the group Gilbert flagged for the age adjustment, so that figure is a within age standing rather than a common scale value.

What did Gilbert say the numbers meant?

Very little. He wrote that they merely confirmed that the most successful men in any sphere of activity, whether politics, industry, militarism or crime, tend to be above average, and that the IQ indicates nothing but the mechanical efficiency of the mind and has nothing to do with character or morals.

Which subtests were in the battery?

Eight. Memory span for number series, arithmetic of increasing difficulty, common sense questions and verbal similarities on the verbal side; a substitution code test, object assembly, block designs and picture completion on the performance side.

Why does removing vocabulary and information matter?

Those two subtests are the most saturated with a specific language and culture and also among the strongest indicators of crystallized verbal ability. Removing them shifts the composite toward memory and constructional tasks and away from accumulated verbal knowledge.

Can these figures be compared to a modern IQ score?

No. The reference sample was American and from the 1930s, the instrument was translated and abridged by the examiner, the item records are unavailable, and no correction factor exists that would convert the figures onto any current scale.

Does the Flynn effect let you adjust the numbers?

No. Pietschnig and Voracek reported annual full scale gains of 0.28 IQ points across 271 samples and almost 4 million participants in 2015, but a meta-analytic trend describes average movement across many samples and is not a conversion factor for one administration.

Why does the Hess figure carry an asterisk?

Gilbert's footnote states it is given on the basis of a retest after recovery. Kelley, writing the same year, reported the range differently, giving 115 to 120 from the Rorschach and a confirming intelligence test.

Did the two examiners agree about what the testing showed?

No, and they broke over it. Kelley concluded the defendants were not a distinct psychological type and that similar personalities could be found anywhere. Gilbert argued they shared a describable pathology and spent his career on that position.

What did the Rorschach reanalyses find?

Harrower reported in 1976 that ten experts assessing the records blind against matched controls did not identify the Nazi protocols as such. Zillmer, Archer and Castino concluded in 1989 that the records cannot be grouped into one specific mental disorder.

Did anyone argue the opposite?

Yes. Miale and Selzer's 1975 volume, which Gilbert introduced and supplied records for, argued for a distinctive pathology, and Resnick and Nunno attempted an actuarial demonstration in 1991. Dimsdale wrote in 2015 that controversy over the records remains.

Did Kelley write anything about Hitler himself?

Yes, and it illustrates the problem. He published a psychiatric characterization of a man he never met, examined or tested, built from documents and interviews with associates. It is a documented opinion, not a finding.

Does a high test score say anything good about a person?

No. A score describes performance on standardized tasks relative to a reference group and carries no moral content in either direction. Gilbert stated that explicitly on the page that carried his own results.

Why are the Nuremberg records still studied?

Because the program set out to test whether measurable psychological characteristics distinguished the defendants from other people. The reanalysis literature has largely converged on the answer that they did not, and that is the durable finding of the episode.

Are the primary sources available to read?

Yes. Gilbert's Nuremberg Diary, Kelley's 22 Cells in Nuremberg and Miale and Selzer's The Nuremberg Mind are all scanned and publicly readable, and the reanalysis papers are indexed with DOIs.

What would a defensible score have required?

A named instrument with a published manual, a date and examiner with recorded conditions, a reference sample the raw performance was converted against, and a document a third party can open. The Nuremberg table has all four in adapted form. The circulating Hitler figures have none.

Take the assessment

You get a profile, not a number

ACIS measures six CHC domains across 20 subtests and reports each one with its own normed score and confidence interval, so you can see where you are strong and where you are not.

Free trial, no card required. Full report from $15.