The Pygmalion effect and IQ: the 1968 claim, the critiques, and what survives
In 1968 Robert Rosenthal and Lenore Jacobson reported that children whose teachers expected them to bloom gained more IQ points, and the claim reached newspapers and textbooks. This page reads the book's own tables, the test behind them, the critiques that followed and the later synthesis of 18 experiments, then states what survives: a small, conditional effect of expectations on performance, and an effect on IQ that is zero or very small, depending on who reads the data.
In the 1968 Oak School experiment, the children named to teachers as likely bloomers were chosen by a table of random numbers, not by any test result.
0 The short answer
The Pygmalion effect is real as a small effect of teacher expectations on student performance, but the 1968 claim that expectations raised children's IQ by large amounts did not survive scrutiny. Rosenthal and Jacobson reported that first graders named as likely bloomers gained 15.4 IQ points more than classmates, a difference based on 7 children, and later reviewers disputed both the test scores and the analysis. Raudenbush's synthesis of 18 experiments found a mean effect size of 0.11, and an effect near zero in the studies where teachers had known their pupils for more than two weeks. Jussim and Harber concluded in 2005 that the effect of teacher expectations on IQ ranges from nonexistent to small.
15.4
IQ points: the first grade difference in mean gains printed in the 1968 book, with 7 children in the experimental group.
0.11
Mean effect size in Raudenbush's 1984 synthesis of 18 experiments, with single studies running from 0.55 down to minus 0.13.
2 weeks
Teacher contact beyond which the same synthesis found no average effect (a mean of minus 0.04 across eight studies).
1 What Is the Pygmalion Effect, and Where Does the Name Come From?
The Pygmalion effect is the name for two claims that travel together: that one person's expectations can change how another person is treated and how that person performs, and that the same mechanism can raise intelligence test scores. The label comes from the 1968 book Pygmalion in the Classroom, by Robert Rosenthal of Harvard and Lenore Jacobson, a school principal, published in New York by Holt, Rinehart and Winston. Its closing section, Shaw's Summary, opens with a passage from George Bernard Shaw's play Pygmalion about how a person is treated, and the book's last pages suggest that perhaps the role of the teacher is Pygmalion's role. The underlying idea, the self-fulfilling prophecy, is older. Spitz's 1999 history of the controversy credits Merton's 1948 essay with giving the concept its name.
The classroom study grew out of Rosenthal's earlier work on experimenter bias. In a 1963 experiment with Fode, published in the journal Behavioral Science, student experimenters were told that their laboratory rats came from strains bred for maze brightness or for maze dullness. By the book's account, sixty ordinary rats were divided among twelve experimenters, and the rats believed to be bright learned the task better than the rats believed to be dull. The experimenters who expected better performance also said that they handled their animals more, and more gently. The book then drew the analogy that started the classroom study: if animals could become brighter when their experimenters expected it, children might become brighter when their teachers expected it. Spitz points out that Rosenthal himself ascribed the rat result to unwitting differences in how experimenters treated the animals, and that the rats could hardly have gained suddenly in rat intelligence. Spitz also notes that most intelligence researchers regard general intelligence as a stubborn trait from childhood onward, shaped by gene and environment interactions, which helps explain why a claim of large gains from a teacher's expectation met skepticism.
That distinction between a change in treatment and a change in ability organizes the rest of this page. A teacher who calls on a pupil more often or presents more material can change what the pupil does on a task. Whether that changes the capacity an IQ test is meant to sample is a separate question, treated from the measurement side on the page on what an IQ test measures and the page on what intelligence is. An IQ score records performance on one occasion, and a change in it can have several causes, of which a change in ability is only one.
The first report was a four page article in 1966, Teachers' Expectancies: Determinants of Pupils' IQ Gains, in Psychological Reports. Its abstract states that children reported to teachers as showing unusual potential, who had actually been selected at random, gained significantly more IQ than the other children eight months later, and that the effect operated primarily among the younger children. The book followed in 1968. According to Spitz, The New York Times ran a front page headline on the results in 1967, and the book was reissued in 1992 with its text unchanged.
2 What Did Rosenthal and Jacobson Do at Oak School in 1964?
The Oak School experiment was a randomized field experiment in which every child took a group intelligence test under a false name, and teachers were then told that a randomly chosen fifth of their pupils were about to bloom. According to the book, the school was a public elementary school, called Oak School in print, with about 650 pupils drawn mostly from a lower class community, and about one sixth of its pupils were Mexican children, the only minority group enrolled. Each grade had three classes, tracked mainly by reading performance and called the fast, medium and slow tracks.
In the spring of 1964 the classroom teachers administered the test to every child who might return the following fall, which meant kindergarten through grade five, because the sixth graders were leaving for junior high school. The teachers were told that it was the Harvard Test of Inflected Acquisition, a new test tied to Harvard and the National Science Foundation that predicted which children were about to show a spurt of academic progress. It was in fact Flanagan's Tests of General Ability, a standardized group test that the book calls relatively nonverbal. The disguise gave the authors a pretest of every child and a believable basis for creating favorable expectations.
At the end of the summer, before the teachers met their new classes, each of 18 teachers (one class per grade in each of the three tracks) received a sheet listing from one to nine children said to be among the top 20 percent on the Harvard test. The names, about 20 percent of the school, had in fact been chosen with a table of random numbers. Teachers were told only that they might find it of interest to know which children were about to bloom, and they were asked not to discuss the list with pupils or parents. The book summarizes the design in a sentence: the only difference between the children earmarked for intellectual growth and the others was in the mind of the teacher.
The children were retested with the same test at the end of the first semester in January 1965, at the end of the school year in May 1965, and a final time in May 1966, when they had moved on to new teachers who, the book states, did not know which children had been designated. The teachers administered the first two retests but did not score them. Research assistants who did not know which children belonged to which group scored every test twice, independently. The May 1965 retest, eight months after the lists were handed out and one year after the pretest, was the basic post-test on which the main comparison rested.
The sample shrank as the study ran. Elashoff and Snow counted 478 children who took the pretest, of whom 382, or 80 percent, had at least one retest. The one year comparison in the book covers 255 control children and 65 experimental children. Random assignment with a pretest, blind scoring and an unannounced follow-up are real strengths, and Jussim and Harber, whose 2005 review is discussed below, call it a simple and elegant study. Its weak point, in Raudenbush's phrase its Achilles' heel, is the expectancy induction: researchers can hand a teacher false information, but an effect can follow only if the teacher believes it and acts on it.
3 What Did the 1968 Book Report About IQ Gains?
The book reports that the designated children gained about 4 IQ points more than the control children over one year, that the advantage sat almost entirely in the first and second grades, and that it did not appear in most of the upper grades. In the book's own wording, the control children gained over eight IQ points in the year of the experiment while the designated children gained over twelve. The printed figures are 8.42 and 12.22, a difference of 3.80 points with a one tailed probability of .02. Table 7-1 of the book breaks the year down by grade, and the grade pattern is the part that matters.
Grade in 1964 to 1965
Control children
Control mean gain
Designated children
Designated mean gain
Difference printed in the book
1
48
12.0
7
27.4
15.4 (one tailed p .002)
2
47
7.0
12
16.5
9.5 (p .02)
3
40
5.0
14
5.0
0.0
4
49
2.2
12
5.6
3.4
5
26
17.5
9
17.4
0.0
6
45
10.7
11
10.0
0.7 in favor of the controls
All grades
255
8.42
65
12.22
3.80 (p .02)
The table reproduces Table 7-1 of Rosenthal and Jacobson (1968), with the differences as the book prints them. The 15.4 point first grade gap, the largest in the book, compares a mean gain of 27.4 points in 7 designated children with a mean gain of 12.0 in 48 controls. The second grade gap of 9.5 points rests on 12 designated children. Together those two groups are 19 of the 65 designated children, about 29 percent, which is our arithmetic on the book's counts. Elashoff and Snow make the same point in their reanalysis: the significant overall advantage rests on the 19 first and second grade children in the experimental group. The book itself reports that the grade by treatment interaction reached only the .07 level.
The book also reports the share of children who gained a given amount. Among the 95 control and 19 designated children of grades one and two, 49 and 79 percent gained at least 10 IQ points, 19 and 47 percent gained at least 20, and 5 and 21 percent gained at least 30. The test yields separate verbal and reasoning scores. In grades one and two, verbal gains were 4.5 points for controls and 14.5 for designated children, a difference of 10.0 (p .02), while in grades three to six controls gained 9.6 and designated children 8.0. In reasoning IQ the whole school gained 15.73 points among controls and 22.86 among designated children, a difference of 7.13 (p .005), and in grades one and two the gains were 27.0 and 39.6, a difference the book prints as 12.7 (p .03).
Two features of those numbers deserve attention. The control children in grades one and two gained 27.0 reasoning IQ points with no special expectation attached, which the authors called remarkable, speculating that taking part in an experiment may itself benefit children. The second feature is the distance between the tables and the prose. Elashoff and Snow list statements in which the book generalizes from the tables, and the book's summary chapter says that a change in teacher expectation can lead to improved intellectual performance, but the tables show an advantage in two grades and not in the other four.
The two year follow-up in May 1966 adds a second caution. The book's Table 9-6 gives a mean total IQ gain of 4.63 points for 196 control children and 7.30 for 47 designated children, a difference of 2.67 whose probability is printed in parentheses as .13, which is not significant. Only the original fifth graders showed a clear advantage, 11.1 points (p .01), although they had shown none a year earlier, which the book calls a baffling question. Spitz adds that among the lowest two grades combined, controls averaged a total IQ of 101 at the basic post-test and the same at the final testing, while designated children fell from 117 to 109.
4 Which Test Measured the Gains, and What Does Its Score Mean?
The IQ in Pygmalion in the Classroom was a ratio score from a group test, converted from raw points through mental age tables, and its behavior at the youngest ages and at the extremes is where most of the controversy lives. The instrument was Flanagan's Tests of General Ability, known as the TOGA, in its 1960 edition. The book describes three elementary forms, for kindergarten to grade 2, grades 2 to 4 and grades 4 to 6, each with a verbal and a reasoning subtest added to give a total. According to Spitz, every item was pictorial and multiple choice with five options. The book says the verbal items were read aloud by the teacher, while the reasoning items were self-administered and timed. Describing formats only: the verbal subtest used picture items on information, vocabulary and concepts, and the reasoning subtest asked children to find the drawing that differed from four others. In the vocabulary of the CHC model, our reading is that the reasoning format samples fluid reasoning (Gf) and the verbal format samples crystallized knowledge (Gc). The page on nonverbal IQ tests covers tests built on figures rather than words, and the page on why IQ tests are timed covers what a clock adds to a score.
Raw points became an IQ in two steps. Spitz reports that the TOGA manual converts raw scores to mental ages, with age equivalents running from 0.5 to 16.5 years, and that the IQ was then mental age divided by chronological age, times 100. That is a ratio IQ, the older method described on the page on mental age. The page on the history of IQ testing places that method in context. Modern batteries report a deviation score, a position among same age peers on a scale with a mean of 100 and a standard deviation of 15, as the page on how IQ is calculated and the page on how IQ scores are normed explain. As our own arithmetic, a six year old one year ahead scores 117 and a twelve year old one year ahead scores 108, so a ratio score gives larger numbers to the same advantage in younger children. Spitz adds that the TOGA norms did not support IQs below 60 or above 160, so IQs outside that range in the Oak School data came from inappropriate extrapolation.
The guessing floor is the most concrete problem. Elashoff and Snow note that the kindergarten to grade 2 form has 63 items with five choices each, so a child who marked answers at random would be expected to get about 13 right. Their conversion shows that a raw score of 8 for a six year old gave an IQ of 50, a raw score of 13 gave 67 and a raw score of 20 gave 83. Across that stretch, 12 more correct answers raised the IQ by 33 points, about 2.75 IQ points per answer, which is our arithmetic on their figures. A child who attempted more items at the post-test could therefore gain many IQ points without a change in anything the test was designed to sample. Rosenthal himself said that the very low pretest IQs were earned because many children attempted few items, as Spitz reports. Wineburg argued in 1987, according to Spitz, that more attempted items could just as well reflect misunderstood instructions, uncontrolled administration, teacher coaching, encouragement to guess, or chance.
Who administered the test is the second concern. Classroom teachers gave the tests, and the TOGA directions left the person administering some freedom, according to Spitz, who adds that group tests are in general more open to artifacts than individual tests. Rosenthal and Jacobson defended group testing on logistical grounds and as a safeguard against examiner expectancy, and they had research assistants score the answer sheets without knowing group membership. Blind scoring is a real safeguard, but the concern is about what happened during administration, such as encouraging children to keep going, which a blind scorer cannot detect afterward. The page on reliability and validity explains why a score depends on conditions of administration as well as on the items.
5 What Did Thorndike and Snow Object To in 1968 and 1969?
Within a year of publication, two reviewers argued that the problem lay in the test scores themselves, not only in the analysis, because scores for the youngest children were too low, too variable and too sensitive to test taking behavior to carry the conclusion. Robert L. Thorndike's review in the American Educational Research Journal in 1968 was blunt. As Spitz reports it, Thorndike judged the book so defective technically that he regretted it had left the original investigators. His central evidence was the pretest. One classroom of 19 children about to enter first grade had a mean reasoning IQ of 31, and the 63 children entering first grade had a mean reasoning IQ of 58. Thorndike estimated that scores that low needed a raw score of only about 2, below what random answering would produce. At the other end, he questioned a post-test mean reasoning IQ of 150 for six designated children in a fast track second grade class, which by his calculation would have required perfect scores, and he asked how such scores could have a standard deviation near 40.
Rosenthal replied in 1969, and Spitz summarizes the reply. The very low IQs were earned because many children attempted very few items. Reasoning scores predicted, at an above chance level, the track that kindergarten teachers recommended, and a year later reasoning IQ correlated .49 with first grade teachers' assessments of future success. Reasoning IQ gains favored the designated children in 15 of 17 classrooms. The mean of 150 followed from mental ages of 16.5, 16.5, 10, 10, 10 and 8.9 years, about 1.5 times the group's mean chronological age. Thorndike's short rejoinder, again as Spitz reports it, held that age equivalents are an unsatisfactory equal unit scale when extrapolated far beyond the ages at which a test was standardized, and that a teacher who encouraged pupils to guess at a few more items could produce a measurable gain through ordinary luck. Rosenthal pointed to predictive correlations and to consistency across classrooms. Thorndike pointed to what the scores at the extremes could mean at all.
Richard Snow's review, Unfinished Pygmalion, appeared in Contemporary Psychology in April 1969 after he had requested and received the raw scores from Rosenthal, according to Spitz. Snow wrote that the study suffers from serious measurement problems and inadequate data analysis, as Jussim and Harber quote him. He said that the TOGA lacked adequate norms for the youngest children, especially those from lower socioeconomic backgrounds, and repeated the implausible first grade pretest means of 31, 47 and 54 in different classes. He also listed individual children, among them one with a pretest reasoning IQ of 17 and later scores of 148, 110 and 112, and one whose verbal IQs across four testings were 183, 166, 221 and 168. By Spitz's count, seven of the 12 scores he listed for three children fell outside the TOGA norm range of 60 to 160. Arthur Jensen added in 1969 that teachers should not have administered the group test, and Spitz reports that Rosenthal answered that the authors had in fact compared classrooms, with larger effects.
Jussim and Harber answer one version of the criticism. Random unreliability makes group differences harder to find, so a difference found despite it is not made less credible by it. That is correct as far as it goes. The critics' strongest points were of another kind. Scores below the level of chance, IQs extrapolated past the norms and behavior that a teacher can influence at the post-test are systematic problems, not random noise, and they can create an apparent gain. The distinction between noise and bias is our reading, not a quotation from either side, and it matters because the usual defense of an unreliable test does not reach a biased one.
6 What Did Elashoff and Snow Find When They Reanalyzed the Data?
The 1971 book Pygmalion Reconsidered, by Janet Elashoff and Richard Snow of Stanford, is the most detailed critique, and its conclusion was that the study neither demonstrated an expectancy effect nor ruled one out in the first two grades.Pygmalion Reconsidered was published by Charles A. Jones Publishing and reanalyzed the scores Rosenthal supplied. Rosenthal and Rubin replied in the same volume, in a chapter titled Pygmalion Reaffirmed, so the book is also a record of the exchange. Elashoff and Snow's criticisms fall into four groups, and the first is reporting. They found that text and tables disagreed, that conclusions were overdramatized, and that labels presupposed their own interpretation: a simple pretest to post-test difference was called intellectual growth, and the difference between two groups was called an expectancy advantage even though it is not always positive. Every analysis was done on group means, yet conclusions were written about individual children, as in the statement that the entire school benefited.
The second group concerns sampling. Of 478 children pretested, 382 were present for at least one post-test, a loss of 20 percent. Elashoff and Snow quote the book's own remark that children who move in and out of the school seldom belong to the high achieving third, and conclude that the remaining children cannot be treated as a random sample. Among those who remained, the pretest scores of the experimental group were consistently higher than those of the controls, despite random assignment. Rosenthal and Jacobson checked whether brighter children had gained more by correlating pretest scores with gains, and Elashoff and Snow showed that this check cannot work, because gain scores are expected to correlate negatively with pretest scores. The third group concerns measurement, covered above: many IQs below 60 or above 160 that the authors did not discuss, and striking sequences for single children, such as total IQs of 55, 102, 95 and 104 across four testings, or reasoning IQs of 0, 77, 82 and 143.
The fourth group is the teachers. After the study, the book reports, teachers could not recall accurately, nor even choose accurately from a longer list, which of their pupils had been designated. Elashoff and Snow treated this as a puzzle for the theory: if the effect existed, it must operate subtly and without conscious awareness on the teachers' part. That reading is possible, and it is also consistent with there being little effect to explain.
The reanalysis itself reached a cautious verdict. Elashoff and Snow found no treatment effect in grades 3 through 6. For grades 1 and 2 they wrote that the experimental and control groups differed greatly on the pretest, so that the data cannot give clear conclusions. They saw enough suggestion of an effect there to warrant further research, and they said the experiment certainly does not demonstrate an expectancy effect or indicate its size. That is an argument about the strength of the original evidence, not a claim that no teacher expectation effect exists, and Snow himself wrote in 1995, as quoted by Spitz, that teacher expectancies can influence classroom teaching and learning at least sometimes. The dispute was over IQ and over a single study, and the next section turns to data from many studies.
Raudenbush's synthesis of 18 experiments is the central quantitative test of whether induced teacher expectations change pupil IQ, and its main result is that the effect depended on how long teachers had known the children before researchers gave them the information. Raudenbush reports that Baker and Crist, reviewing 25 early studies in 1971, found that none of the nine studies with IQ as the outcome showed a significant overall effect. Raudenbush's 1984 paper pooled 18 experiments, all with IQ as the outcome and children in grades 1 to 7, the Oak School study among them. He excluded three studies that involved adult learners, children with intellectual disability or an ambiguous outcome measure. His effect size d is the treatment effect in IQ points divided by the control group's standard deviation at the post-test.
The 18 effect sizes ranged from 0.55 to minus 0.13, with a mean of 0.11 and a standard deviation of 0.20. Five studies reached statistical significance, three at the .05 level and two at the .01 level, and none was significantly negative. The Oak School study had d = 0.21 (one tailed p = .016) and one week of prior teacher contact. Three of four methods for combining probabilities rejected the null hypothesis of no effect, and the method weighted by sample size did not. Raudenbush noted that the pooled mean resembles the 0.16 that Smith found for 22 IQ estimates in 1980, against 0.38 for 78 estimates of effects on achievement.
The pooled mean hides the finding that matters. Raudenbush predicted that the more weeks teachers had spent with their pupils before the induction, the smaller the effect would be, and prior contact in the 18 studies ranged from 0 to 24 weeks. The correlation between effect size and weeks of contact was minus .55, and after a transformation to straighten a curved relationship it was minus .77, so that weeks of prior contact accounted for 59 percent of the variability in effect sizes. The table gives his summary by weeks, with the binomial effect size display he printed beside it, which shows how many experimental and control children would qualify for a higher academic track if the median IQ were the cutoff.
Weeks of teacher and pupil contact before the induction
Studies
Mean effect size d
Experimental children above the cutoff (percent)
Control children above the cutoff (percent)
0
4
0.32
58
42
1
3
0.26
56.5
43.5
2
3
0.08
52
48
More than 2
8
minus 0.04
49
51
Source: Raudenbush (1984), Tables 7 and 8. The paper prints a mean of minus 0.04 for the eight high contact studies in its text and Table 7, and minus 0.06 in the note to Table 3.
Split at two weeks, ten studies with 2 weeks or less of contact had a mean d of 0.23 and each of four combined significance tests found an effect greater than zero, while the eight studies with more than two weeks found none in either direction. Whether the test was given to a group or to an individual, and whether the person giving it knew which children were designated, made no difference among low contact studies. Grade mattered in a pattern the original study did not predict: among low contact comparisons the mean was 0.31 in grades 1 and 2 and 0.04 in grades 3 to 6, and it reappeared at 0.25 in grade 7, where researchers had taken care to keep teachers from knowing their pupils beforehand.
Raudenbush's explanation is about belief. A teacher who already knows a child from weeks of contact is likely to reject false information about that child, so the treatment was never delivered. He also noted a limit: the experiments test artificially induced expectations, which leaves open the effects of naturally occurring ones. An experiment in which the induction fails cannot show that expectations do not matter.
8 How Do Supporters and Critics Read the Synthesis?
The synthesis supports two readings, which is why the dispute outlasted it: supporters see a real effect whenever the expectation was credible, and critics see a few strong results in a set that averages close to nothing. Rosenthal's 1997 address, Interpersonal Expectancy Effects: A Forty Year Perspective, cites Raudenbush as finding very strong evidence, with a correlation of .67, that substantial teacher expectancy effects on IQ appear only when the induction is credible. Spitz points out that .67 was Rosenthal's own figure from a split at two weeks, while Raudenbush printed minus .77 for the continuous measure.
The critics' objections are specific. Snow's 1995 reply, as Jussim and Harber summarize it, held that some studies produced reversals and that the median effect size of .035 was a better estimate than the mean of .11. Spitz adds that Snow questioned counting a condition in which tutors were familiarized with the test, which produced the largest effect size in the set, 0.85, and which is why 18 studies yield 19 effect sizes. Snow also reanalyzed the Oak School data, according to Jussim and Harber, and reported that five children with gains averaging over 90 IQ points accounted for the group difference, and that excluding all scores outside the TOGA's range of 60 to 160 removed the effect.
Raudenbush answered with a reanalysis in 1994 using random effects models. Jussim and Harber report that the effect size for the four studies with no prior teacher and pupil contact was .43, a correlation of about .2, and that the remaining 14 studies still showed no overall effect. Spitz reports that the estimated effect fell by .17 for each added week of contact up to two weeks. The studies with no prior contact show an effect, and the studies in which teachers had known their pupils for weeks show none.
The verdicts differ in tone more than in content. Spitz concluded that the data behind Raudenbush's meta-analyses are not strong enough to support the blanket claim that teacher expectations affect pupils' intelligence. Jussim and Harber, whose 2005 review in Personality and Social Psychology Review covers 35 years of research, wrote that self-fulfilling effects on IQ range from nonexistent, on the critics' reading, to small, on Rosenthal's and Raudenbush's, and that what is certain is that the hypothesis of large and dramatic effects on IQ has been disconfirmed. The factual disagreement is small, between zero and very small. Our reading is that both camps agree on the practical point: in a real classroom, where teachers know their pupils from the first weeks, the conditions under which the effect appeared in experiments rarely hold, and the effect on IQ scores that remains is not large.
9 How Large Are Teacher Expectation Effects, in Standard Deviations and in People?
Across the literatures, teacher expectation effects are small by conventional standards, and how large they look depends on whether they are stated as a standard deviation, a correlation or a share of students. Jussim and Harber put the Oak School result itself at an effect size of .30, a correlation of .15 and a mean difference of about 4 IQ points, and note that effect sizes of .30 or less are conventionally considered small. Significant effects appeared in two of six grades in the first year and one of five in the second, so the prophecy did not occur in 8 of 11 grade comparisons. For teacher expectations in general they conclude that effects average a correlation of about .1 to .2, which in Rosenthal's binomial display means that 55 to 60 percent of students with high expectations end up above average against 45 to 40 percent with low expectations, a change in achievement for about 5 to 10 percent of students.
Rosenthal's own summaries report larger figures and pool far more kinds of study. His 1997 address reports a mean effect size of d = .62, or r = .30, across 479 studies of interpersonal expectancy effects, and, after Rosenthal and Rubin's 1978 review of the first 345 studies, d = .54 and r = .26 for the learning and ability domain, whose examples include IQ test scores and verbal conditioning. Spitz objects that merging expectancy and intelligence studies with other expectancy studies, as those reviews do, dissolved the problem of replicating Pygmalion itself. A pooled figure across rats, inkblots and reaction times does not say how large the effect is on a child's IQ score.
Source
Design
Measure
Size as printed in the source
Rosenthal and Jacobson (1968)
Randomized field experiment, 18 classrooms
Total IQ gain on the TOGA
3.80 points overall; 15.4 in grade 1; d .30 and r .15 per Jussim and Harber
Raudenbush (1984)
Meta-analysis, 18 experiments
Pupil IQ
Mean d = 0.11; 0.32 with no prior contact; minus 0.04 beyond two weeks
Raudenbush (1994), per Jussim and Harber
Random effects reanalysis
Pupil IQ
d = .43 for four studies with no prior contact
Rosenthal and Rubin (1978), per Rosenthal (1997)
Review of 345 experiments in eight domains
Learning and ability, including IQ test scores
d = .54, r = .26
Rosenthal (1997)
479 studies of interpersonal expectancy
Mixed outcomes
d = .62, r = .30
Jussim and Harber (2005)
Review of 35 years of research
Student achievement
Typical r = .1 to .2
McNatt (2000)
Meta-analysis, 17 workplace studies
Performance
Average d = 1.13
Kierein and Gold (2000)
Meta-analysis, 13 workplace effect sizes
Performance
Overall d = 0.81
A last point separates prediction from causation. Jussim and Harber report that teachers' expectations predict student achievement mostly because they are accurate and not because they are self-fulfilling. By their estimate, about 75 percent of the predictive validity of teacher expectations for standardized test scores reflects accuracy and about 25 percent reflects self-fulfilling prophecy. A teacher who expects a pupil to do well on a test, and is right, is usually predicting an outcome that the pupil's existing ability already made likely.
10 Does the Effect Carry Into Workplaces, and What Is the Golem Effect?
Reviews of workplace studies report larger effects than the classroom IQ literature, but they measure performance on tasks and not intelligence, and they report the conditions under which the effect is strongest. McNatt's meta-analysis in the Journal of Applied Psychology pooled 17 studies of Pygmalion interventions with adults in management contexts, 58 effect sizes and 2,874 participants, with outcomes such as exam scores, performance appraisals and physical output. The average was d = 1.13, with wide variation, and results were stronger in the military, with men, and for people about whom low expectations were initially held. Kierein and Gold's meta-analysis of work organizations found an overall d of 0.81 across 13 effect sizes, stronger when initial performance was low and in military rather than business settings. The page on collective intelligence takes up how performance in groups relates to the abilities of their members.
The Golem effect is the mirror image, poorer performance resulting from low or negative expectations, and the name comes from the title of Babad, Inbar and Rosenthal's 1982 paper, Pygmalion, Galatea, and the Golem. Wang, Rubie-Davies and Meissel describe Babad's later work as finding such effects only in the classrooms of highly biased teachers. Jussim and Harber, however, found little and contradictory evidence on whether negative expectations are more powerful than positive ones, and the Oak School study could not test it, because the authors induced only favorable expectations. Wang and colleagues report mixed findings on whether students from stigmatized groups are more susceptible to teacher expectations.
11 How Could Expectations Change Performance Without Changing Intelligence?
The mechanisms proposed by Rosenthal and later reviewers all run through what teachers and students do, which can change performance on a task or a test without changing the ability the test was built to sample. Rosenthal's 1997 summary of the four factor theory, drawing on Harris and Rosenthal's 31 meta-analyses, reports that teachers appear to create a warmer climate for students they expect to do well (a mean correlation of .29), to teach them more and harder material, called input (.27), to give them more chances to respond, called output (.17), and to give them more differentiated feedback (.10). Those are correlations between expectation and teacher behavior, and between behavior and student response. They are not effects on IQ.
The 1968 book added one study at press time. In Beez's experiment with sixty preschoolers in a summer Head Start program, each child was taught a series of symbols by one teacher. Teachers who had been led to expect good learning tried to teach eight or more symbols in 87 percent of cases, against 13 percent for teachers led to expect poor learning, and 77 percent of the children said to have better prospects learned five or more symbols, against 13 percent. The authors report that the difference remained, at about half its size, when teaching effort was controlled. The outcome was learning of taught material, judged by an experimenter who did not know the expectation, and it was not a score on an intelligence test.
Wang, Rubie-Davies and Meissel's systematic review of 144 studies from 1989 to 2018 adds two cautions. Student self-concept and self-efficacy appear to mediate expectation effects on achievement, and nearly 40 percent of the studies that related expectations to achievement did not control for students' baseline achievement, so some of the association may reflect where students started. A pupil who is encouraged to keep working may attempt more items, concentrate longer or become more familiar with the format, and each of those changes a score. The page on how to prepare for an IQ test separates preparation that helps a person from preparation that distorts a score. Whether such a change reflects the ability itself is a separate question that these designs were not built to answer. The page on intelligence and the brain covers what imaging research says about the biology of ability.
12 What Does the Evidence Support, and How Should a Change in an IQ Score Be Read?
The evidence supports a small, conditional effect of teacher expectations on performance, it does not support the 1968 claim of large IQ gains, and the Standards for Educational and Psychological Testing say how any change in a score should be read. Stated narrowly, five things hold. The Oak School randomized experiment found a 3.80 point average advantage, carried by 19 children in grades 1 and 2 and not significant at the two year follow-up. The best synthesis of IQ experiments found a mean d of 0.11 that depended on weeks of prior contact and was near zero beyond two weeks. Effects on achievement across the wider literature are small, typically r = .1 to .2. Workplace studies find larger effects on performance. And teacher expectations predict outcomes mainly because they are accurate. Every figure above is a group average, so the evidence does not say what any one child or adult would gain, it does not show that a teacher's label raises an individual's IQ, and it does not show that a low score was caused by a low expectation. ACIS norms cover ages 16 to 90, so nothing here is an offer to test children.
The 2014 Standards for Educational and Psychological Testing, published jointly by AERA, APA and NCME, speak directly to the Oak School data. Standard 2.4 asks that when an interpretation emphasizes differences between two observed scores, reliability and precision data, including standard errors, be provided for the differences, and its comment notes that the reliability of change scores can be much lower than that of the separate scores. Standard 13.2 asks that when change or gain scores are used, their construction, technical qualities and limitations, and the time between administrations be reported, with care to avoid practice effects. Standard 6.1 asks administrators to follow the developer's standardized procedures. The APA Guidelines for Psychological Assessment and Evaluation, approved in March 2020, add in Guideline 8 that normative data may not be accurate when administration or scoring departs from standardization or when floor and ceiling effects apply, and that adhering to standardized conditions minimizes confounds that lead to misinterpretation. Applying them to Oak School, teacher administered group tests, IQs extrapolated beyond the norms and change scores from a guessing floor are each a point these documents address. That application is ours and is not a finding of either document.
An ACIS report gives percentiles and a 95 percent confidence interval, and the page on how an IQ score differs from a percentile explains why the two cannot be averaged. A reader can see that one number stands for a range of likely values, and a difference between two sittings carries more error than either score alone. The page on how IQ changes with age, the page on what training can move and the page on the Flynn effect, a different named effect that concerns norms across generations, take the question of change further. The technical manual covers the instrument.
Every figure above is traceable to one of the following, and each is linked at the point where it is used. Where a number is our own arithmetic on a published figure, the text says so. The reviews by Thorndike (1968) and Snow (1969) are listed by their verified bibliographic records, and what they said is reported as Spitz (1999) and Jussim and Harber (2005) summarize and quote them. The same applies to Smith (1980) and Raudenbush (1994), which are reported through Raudenbush (1984) and Jussim and Harber (2005). The full texts of the Standards, the APA guidelines, the two books and the papers by Raudenbush, Spitz, Jussim and Harber, Wang and colleagues and Rosenthal (1997) were read on October 6, 2026. For Rosenthal and Jacobson (1966), Rosenthal and Rubin (1978), McNatt, and Kierein and Gold, the abstracts were read, and for Thorndike, Snow, and Babad and colleagues only the verified bibliographic records were checked.
American Educational Research Association, American Psychological Association and National Council on Measurement in Education. Standards for Educational and Psychological Testing. AERA, 2014, Standards 2.4, 6.1, 12.11 and 13.2.
American Psychological Association, APA Task Force on Psychological Assessment and Evaluation Guidelines. APA Guidelines for Psychological Assessment and Evaluation. Approved by the APA Council of Representatives, March 2020, Guidelines 7 and 8, read October 6, 2026.
Babad E, Inbar J and Rosenthal R. Pygmalion, Galatea, and the Golem: Investigations of biased and unbiased teachers. Journal of Educational Psychology, 1982, volume 74, issue 4, pages 459 to 474.
Elashoff J D and Snow R E. Pygmalion Reconsidered: A Case Study in Statistical Inference. Charles A. Jones Publishing, Worthington, Ohio, 1971.
Jussim L and Harber K D. Teacher Expectations and Self-Fulfilling Prophecies: Knowns and Unknowns, Resolved and Unresolved Controversies. Personality and Social Psychology Review, 2005, volume 9, issue 2, pages 131 to 155.
Kierein N and Gold M. Pygmalion in work organizations: a meta-analysis. Journal of Organizational Behavior, 2000, volume 21, issue 8, pages 913 to 928.
McNatt D. Ancient Pygmalion joins contemporary management: A meta-analysis of the result. Journal of Applied Psychology, 2000, volume 85, issue 2, pages 314 to 322.
Raudenbush S W. Magnitude of teacher expectancy effects on pupil IQ as a function of the credibility of expectancy induction: A synthesis of findings from 18 experiments. Journal of Educational Psychology, 1984, volume 76, issue 1, pages 85 to 97.
Rosenthal R and Jacobson L. Teachers' Expectancies: Determinants of Pupils' IQ Gains. Psychological Reports, 1966, volume 19, issue 1, pages 115 to 118.
Rosenthal R and Jacobson L. Pygmalion in the Classroom: Teacher Expectation and Pupils' Intellectual Development. Holt, Rinehart and Winston, New York, 1968.
Rosenthal R and Rubin D B. Interpersonal expectancy effects: the first 345 studies. Behavioral and Brain Sciences, 1978, volume 1, issue 3, pages 377 to 386.
Snow R E. Unfinished Pygmalion. Contemporary Psychology, 1969, volume 14, issue 4, pages 197 to 199.
Spitz H H. Beleaguered Pygmalion: A History of the Controversy Over Claims That Teacher Expectancy Raises Intelligence. Intelligence, 1999, volume 27, issue 3, pages 199 to 234.
Thorndike R L. Review of Pygmalion in the Classroom by Robert Rosenthal and Lenore Jacobson. American Educational Research Journal, 1968, volume 5, issue 4, pages 708 to 711.
Wang S, Rubie-Davies C M and Meissel K. A systematic review of the teacher expectation literature over the past 30 years. Educational Research and Evaluation, 2018, volume 24, issue 3 to 5, pages 124 to 179.
14 Frequently Asked Questions
Is the Pygmalion effect real?
Yes, in a limited form. Teacher expectations have small effects on student performance, with Jussim and Harber reporting typical correlations of about .1 to .2. The stronger claim, that expectations raise IQ substantially, is not supported. IQ experiments average an effect size of about 0.11 and show essentially nothing when teachers already know their pupils.
What is the Pygmalion effect in simple terms?
It is the idea that high expectations from someone in authority, such as a teacher, can lead a person to perform better, while low expectations can lead to worse performance. The name comes from the 1968 book Pygmalion in the Classroom, which claimed that teacher expectations raised pupils' IQ scores.
Can teacher expectations raise IQ?
Not by large amounts, according to the research reviews. Jussim and Harber concluded that effects on IQ range from nonexistent to small, and Raudenbush's synthesis found essentially no average effect when teachers had known pupils for more than two weeks. A raised score is also not proof that ability changed.
What did the Pygmalion in the Classroom study find?
Rosenthal and Jacobson reported that children randomly named as likely bloomers gained about 3.8 IQ points more than classmates over one year, with the gap concentrated in first and second grade. Reviewers later disputed the test scores and the analysis, and the gap at the two year follow-up was not significant.
Was the Pygmalion study replicated?
Many replications were attempted, with mixed results. Raudenbush pooled 18 IQ experiments and found a small mean effect that depended on how long teachers had known the pupils. Spitz judged the evidence too weak to support the claim that expectations raise intelligence. Expectancy effects on other outcomes are better supported.
Is the Pygmalion effect a self-fulfilling prophecy?
It is one type of self-fulfilling prophecy, in which a belief leads people to behave in ways that make it come true. Merton gave the concept its name in 1948, according to Spitz. Rosenthal's contribution was to test the idea experimentally with experimenters and teachers, and later with managers.
Has the Pygmalion effect been debunked?
The large IQ claim has been seriously undermined, but the broader idea has not been debunked. Critics found problems in the original data, and meta-analyses found small effects. Jussim and Harber conclude that classroom self-fulfilling prophecies do occur, though they are typically small and may fade rather than accumulate.
What test did Rosenthal and Jacobson use?
They used Flanagan's Tests of General Ability, a standardized group test that the book calls relatively nonverbal, with verbal and reasoning subtests. Teachers were told it was the Harvard Test of Inflected Acquisition, a new predictor of academic spurts. Classroom teachers administered it, and assistants unaware of group membership scored it.
How many children were in the Oak School study?
About 650 children attended the school. Elashoff and Snow counted 478 pretested children, of whom 382 had at least one retest. The one year comparison in the book covers 255 control and 65 experimental children, and the first and second grade results rest on 19 experimental children.
What did Thorndike say about Pygmalion in the Classroom?
As Spitz reports, Thorndike judged the book technically defective. He pointed to pretest reasoning IQs so low that they implied raw scores below chance, such as a mean of 31 in one class entering first grade, and to extreme post-test scores. He argued that the testing could not support the conclusions drawn.
What did Raudenbush's meta-analysis find?
Across 18 experiments the mean effect size was 0.11, with single studies running from 0.55 to minus 0.13. Effects were larger when teachers had little prior contact with pupils, and the eight studies with more than two weeks of contact averaged about zero. Group or individual testing made no difference.
How large is the Pygmalion effect?
It depends on the measure. For the Oak School IQ gains, Jussim and Harber put the effect at about .30 standard deviations, or a correlation of .15. For teacher expectations generally they report typical correlations of .1 to .2, a change in achievement for perhaps 5 to 10 percent of students.
What is the Golem effect?
The Golem effect is the negative counterpart of the Pygmalion effect, in which low expectations lead to poorer performance. The name comes from a 1982 paper by Babad, Inbar and Rosenthal. Jussim and Harber found the evidence that negative expectations are stronger than positive ones thin and contradictory.
Does the Pygmalion effect work in the workplace?
Meta-analyses report effects on performance in work settings. McNatt found an average d of 1.13 across 17 studies, and Kierein and Gold found 0.81 across 13 effect sizes, both strongest in military settings. These measure task performance, not intelligence, and ongoing for-profit workplaces are less studied.
Why do IQ scores change between two tests?
Scores can differ because of measurement error, practice, effort, testing conditions or a real change in ability. The Standards for Educational and Psychological Testing note that the error in a change score is higher than the error in either original score, so a single difference should be read with caution.
Is a 15 point IQ gain believable?
It needs strong evidence. In the Oak School data, the 15.4 point first grade gap rested on 7 designated children, and scores for the youngest children included values below the norm range of the test. A gain of that size from a change in expectations is not what meta-analyses of IQ experiments found.
Did the control group's IQ rise too?
Yes. Controls gained 8.42 IQ points on average over the year, and first grade controls gained 12.0. In reasoning IQ, controls in grades one and two gained 27.0 points. The authors speculated that taking part in an experiment may itself help children, and practice on a repeated test is another possibility.
Does the Pygmalion effect mean IQ tests are unreliable?
No. The criticisms concerned one group test used with very young children, scores extrapolated beyond its norms, and conditions of administration. They are reasons to check reliability evidence, norm ranges and standard administration. They do not show that every well normed test given under standard conditions fails to measure ability.
Can retaking an IQ test show that my IQ changed?
Only in a limited way. A difference between two sittings combines measurement error, practice, effort and any real change, so one gap is weak evidence of change. The confidence interval around each score shows how much of a gap could be noise, and a small gap sits easily within it.
Can ACIS test children or measure expectation effects?
No. ACIS is a self-administered online test with adult norms for ages 16 to 90, so it does not test young children. It is not a clinical instrument, and its report gives percentiles and a 95 percent confidence interval so that a score is read as a range and not as an exact value.
What should I take away about expectations and intelligence?
Expectations can shape behavior, effort and opportunities, and those can affect performance on tasks and tests. The evidence does not show that they change intelligence itself by a large amount. Read any announced IQ gain by asking about the test, its norms, how it was given and how large the effect was.
Take the assessment
You get a profile, not a number
ACIS measures six CHC domains across 20 subtests and reports each one with its own normed score and confidence interval, so you can see where you are strong and where you are not.