Reporting IQ in research: the eight facts a score hides
A reported IQ is a number plus the instrument, the norms, the mode and the sample that produced it. This page lists the eight facts a methods section should state beside the score, ties each one to the APA JARS-Quant table and the 2014 Standards, explains why norms collected in 2007 and 2008 inflate scores today, and names the errors that make a score unreadable: prorated totals called full IQ and averaged percentiles.
The WAIS-IV technical manual states that its normative data were collected from March 2007 to April 2008, a date a methods paragraph should carry forward beside the score.
0 The short answer
A reported IQ is only interpretable when eight facts travel with it: the instrument and form, the norms and the year they were collected, the score level, the score scale, how the score was derived, the reliability in your own sample, the administration mode, and the exclusions. The APA's JARS-Quant table asks for the measures, the data collection method, reliability for the scores analyzed, the setting and dates, and the exclusion criteria, and the 2014 Standards for Educational and Psychological Testing adds norm dates and the proctoring status. Norms age: a 2014 meta-analysis of 285 studies puts the Flynn effect at 2.31 points per decade, so a score from 2007 to 2008 norms flatters a 2026 sample.
8
Facts a methods section should state beside an IQ, which this page ties to JARS-Quant and the Standards.
2.31
Standard score points per decade, the meta-analytic mean Flynn effect across 285 studies, according to Trahan and colleagues in 2014.
0.41 and 0.21
Annual IQ point gains for fluid and crystallized tests, according to Pietschnig and Voracek in 2015.
1 Why does a research paper need more than the IQ number?
A bare IQ is a conclusion without its method, and a reader cannot judge, repeat or compare it. The same figure of 112 can come from a supervised battery scored against norms collected recently, from a short online form scored against a decades old table, or from a raw count of correct items relabeled as IQ. The three share a digit and almost nothing else. The Standards for Educational and Psychological Testing (AERA, APA and NCME, 2014), published by the American Educational Research Association and available as an open access PDF, frames the problem from both sides. Test documentation should let users decide which test to use, how to administer it and how to interpret its scores (Standard 7.0), and users in turn must study the developer's materials before adopting a test (Standard 9.2). A methods section is where that user duty becomes visible to a reader.
The second framework is the APA's own. The task force report by Appelbaum and colleagues (2018) in American Psychologist set the Journal Article Reporting Standards for quantitative research, JARS-Quant, and its Table 1 applies to any manuscript reporting new data. We read that table in the 2020 copy that SAGE Publications hosts for authors. It does not mention IQ. It asks authors to define all primary and secondary measures, describe the methods used to collect data, provide instrument information including psychometric properties, report reliability for the scores analyzed, describe where and when data were collected, and state inclusion and exclusion criteria. An IQ score is a measure, so every one of those lines applies to it.
We turn those general lines into eight specific facts. The grouping is ours, and the sources behind each row are the ones this page cites section by section.
Fact to state
What it answers
Where the requirement comes from
1. Instrument, edition and form
Which test produced the score
JARS-Quant measures and instrumentation rows; Standard 7.0
2. Norms and the year collected
Whom the score is compared with, and when
Standards 5.8, 5.9 and 5.11
3. Score level
Full Scale composite, index or single subtest
Standard 2.3
4. Score scale
Raw, scaled, standard or percentile
Standards 5.1 and 5.10
5. Derivation
Full battery, substitution, proration or short form
Standards 5.19 and 5.20 and the publisher manual
6. Precision in your sample
Reliability for the scores you analyzed, and the interval
JARS-Quant psychometrics row; Standards 2.3 and 2.19
7. Mode and conditions
Supervised or not, device, setting, dates
ITC guidelines; Standards 1.10, 3.4 and 6.3
8. Exclusions and missing data
Who was dropped, by what rule, how many
JARS-Quant inclusion, exclusion and data diagnostics rows
A disclosureWe sell ACIS, an online cognitive assessment. This page is a reporting guide, not a product comparison, and it applies the same eight facts to ACIS in its last section, where its limits are stated as plainly as its features.
Two limits apply throughout. This page does not tell a researcher which measure to choose, which is the job of the page on measuring IQ in online studies, and it does not cover how to screen careless responders, which the page on online testing data quality handles. It also gives no ethics committee or journal advice: the author guidelines of a journal override anything here.
2 Which instrument, edition and form produced the score?
Name the instrument by publisher, title, edition and form, because a new edition is a different measurement with its own norms. The Standards say so directly for revisions. When test specifications change between versions, the changes should be identified, and users should be told that converted scores for the two versions may not be strictly equivalent, even when statistical linking is used (Standard 5.20). A paper that writes "the WAIS" without an edition has hidden exactly the information the Standard says a reader needs. The pages on the WAIS-IV, the WAIS-5 and what changes between them show how much moves between editions, and the page on WAIS-5 subtests lists which tasks feed which score.
The WAIS-IV technical and interpretive manual gives a small example of why wording matters. Its authors note a change in the terminology for composite scores, so that the verbal and performance composites of earlier editions became the Verbal Comprehension Index and the Perceptual Reasoning Index. A paper that copies "Verbal IQ" from an older edition's literature into a WAIS-IV study has mislabeled the score. Edition-specific names are part of the instrument description.
The form matters as much as the edition. A battery can be given in full, in a subset, in an abbreviated version or in a language adaptation, and each is a different score source. The Spanish-language edition of the WAIS-IV, published in Mexico by Manual Moderno in 2014 from the Pearson 2008 original, credits its standardization coordination to the Faculty of Psychology of the National Autonomous University of Mexico, so its norms are not the norms of the English edition. When a study uses an adapted or translated version, the language and the adaptation norm source belong in the sentence. A battery assembled from an item pool is likewise a new form, and the form has to be described item type by item type.
Three habits make the instrument sentence complete. First, give the publisher and the edition or version number. Second, give the form: full battery, named subset, short form, or item pool with the number of items per type. Third, when a platform or scoring program converted the raw responses, name it and its version, because the conversion is part of the measurement. That third habit is our recommendation rather than a quoted standard, but it follows from the JARS-Quant instrumentation line, which asks for information on the instruments used.
Never reproduce items in a paper or an appendix. Items of commercial batteries are secure test material, and a methods section describes formats (for example, "a matrix completion task with a fixed time limit") and never the stimuli. Describing what an instrument asks, without showing it, keeps the methods section replicable by readers who hold a license.
3 Which norms, and from what year?
State the norm table that converted raw performance to a standard score, who it describes, and when its data were collected, because a standard score is only a position within a reference group. The Standards require that norms refer to clearly described populations, including those with whom users will ordinarily wish to compare their examinees (Standard 5.8). They require that reports of norming studies specify the population sampled, the sampling procedures and participation rates, any weighting, the dates of testing and descriptive statistics, so that users can judge the fit of the norms to local examinees (Standard 5.9). Publishers carry a duty to renorm with sufficient frequency, but the comment to Standard 5.11 is blunt that the user remains responsible for avoiding the inappropriate use of out of date norms.
That makes the norm year a reporting item, not a footnote. The WAIS-IV technical and interpretive manual (Wechsler, 2008, Pearson) states that its normative data were established from a sample collected from March 2007 to April 2008, stratified on age, sex, race and ethnicity, education and region against October 2005 census data. A thesis that tests participants in 2026 with that instrument is applying norms that are about 18 years old, by our arithmetic. The manual itself explains why this matters: it cites Flynn's work for the finding that older norms produce inflated scores, and it says scores should rest on norms that are contemporary and representative. The section on the Flynn effect below quantifies the drift.
A clinical chapter offers a rough grading of recency. Farmer and Floyd, in the fourth edition of Contemporary Intellectual Assessment (Flanagan and McDonough, 2018), cite standards from Floyd and colleagues that norming data are good when some were collected within the past 10 years, adequate within 15 years, and inadequate when all were collected more than 15 years ago. That grading was written for evaluating tests used in intellectual disability assessment, so treat it as a screening heuristic, not a rule for research papers. By that heuristic, the 2007 to 2008 WAIS-IV data crossed the 15 year mark in 2023.
Four norm situations arise in practice, and each needs different wording.
A publisher norm table. Name the table, the age band your participants fall in, and the collection years. Say whether you used the printed table or the publisher scoring software.
A norm table from another country or language. Name it and justify using it for your sample, since a Mexican standardization describes a different reference group than a United States one.
No norms at all. Some measures ship as item pools or research instruments without a population table. Scores there are raw counts or sample-relative scores, and calling them IQ is an error covered below.
Sample-standardized scores. If you convert raw scores to z scores or T scores within your own sample, say so. Such scores describe position inside your sample, so a mean of zero is guaranteed by construction and carries no information about the population.
The page on how IQ scores are normed explains the construction of a norm table in general, and the page on the scoring chain shows how a raw score travels to a standard score.
4 Composite, index or subtest: which score level did you analyze?
Say whether the score is a Full Scale composite, an index or a single subtest, because each has its own content, its own norms and its own reliability. The WAIS-IV technical manual shows the layers. Its five composite scores are the Full Scale IQ and four indices, and it adds the General Ability Index as an additional composite. Each subtest is scaled to a mean of 10 and a standard deviation of 3, and the composites to a mean of 100 and a standard deviation of 15. A Full Scale IQ and a Verbal Comprehension Index can both be written "112" and still describe different things: the first pools many tasks and the second pools a few. The page on the Full Scale IQ explains the composite and the page on the General Ability Index explains why a second global composite exists.
The Standards attach a reporting duty to the choice. For each total score, subscore or combination of scores that is to be interpreted, estimates of reliability and precision should be reported (Standard 2.3), and the comment adds that reporting them only for total scores is not enough when subscores are interpreted, since a total can be consistent while subscores are not. The practical rule is simple. Every level you interpret needs its own precision statement, so a study that analyzes six indices needs six reliabilities, not one.
JARS-Quant adds the opposite duty: define all primary and secondary measures and covariates, including measures collected but not included in the report. In an IQ study this closes a common loophole. If a battery produced a composite and four indices and the paper reports only the index that reached significance, the reader cannot count the comparisons. State which level was the primary measure before the analysis, list the others as secondary, and say where the full set is available.
Naming deserves care. Use the publisher label for the score (Full Scale IQ, Working Memory Index, scaled score) and reserve "IQ" for the global composite. An index score is not an IQ, and a subtest score is not an index. When you translate indices into broad abilities, remember that the index names are publisher labels, and the mapping to the CHC abilities Gf, Gc, Gv, Gwm and Gs is a claim that needs its own citation. The page on the CHC model lists the broad abilities, and the page on what IQ measures separates the construct from the composite. For subtest-level analyses, the page on subtest types shows what each family of tasks asks, and the guidance in this section applies with more force, because single subtests are the least reliable score a battery produces.
A last detail: report age. Norm-referenced scores are age-referenced, and a table of mean standard scores by age band is more useful to a reviewer than a single pooled mean when the sample spans decades. Name the age bands of the norm table and the age range of your sample side by side.
5 Raw, scaled, standard or percentile: which scale is the analysis on?
State the scale of every score you report and analyze, because a raw score, a scaled score, a standard score and a percentile rank are four different quantities. The Standards ask that the characteristics, meaning and intended interpretation of scale scores be explained clearly, together with their limits (Standard 5.1), and the same chapter requires that the procedures for constructing reporting scales be described (Standard 5.2). In a paper, that translates into a one line ladder: raw score, converted by the norm table to a scaled score (mean 10, standard deviation 3), summed and converted to a composite standard score (mean 100, standard deviation 15), then read as a percentile rank. The page on the standard deviation of 15 explains why that scale looks the way it does, and the page on IQ score versus percentile separates the two readings.
The practical rule is to analyze standard scores and to describe with percentiles. Standard scores sit on a scale whose steps are equal in size, so means, standard deviations, correlations and regressions behave as they should. Percentile ranks do not. Equal steps on the standard scale cover very unequal shares of people, as the table shows. The figures assume a normal distribution with a mean of 100 and a standard deviation of 15, and the percentile column is our arithmetic.
Standard score
Percentile rank
Gap from the row above
70
2.3
85
15.9
13.6 points of percentile
100
50.0
34.1 points of percentile
115
84.1
34.1 points of percentile
130
97.7
13.6 points of percentile
Every row is 15 standard score points from the next, yet the percentile gap runs from 13.6 to 34.1. Averaging percentiles therefore produces a number that belongs to no scale. Take two participants, one with a standard score of 100 and one with 130. Their percentile ranks are 50.0 and 97.7, which average to 73.9. A standard score of 109.6 sits at the 73.9th percentile, while the average of the two standard scores, 115, sits at the 84.1st percentile. The two "average" results differ by more than five standard score points, by our arithmetic, and only the second is on a scale that supports averaging.
The Standards address the group version of this problem. In the comment to Standard 5.10, they note that the percentile rank of a school's average score cannot be determined if all that is known is the percentile rank of each student, and that one acceptable procedure is to report the percentile rank of the median group member. For a research paper the equivalent advice is to compute the group mean of standard scores and, if a percentile is wanted for description, report the percentile of that mean standard score or the median percentile rank, and say which one it is. The page on percentile calculation converts a standard score to a percentile, and the page on how rare a given IQ is shows the tails, where percentile steps shrink most.
Raw scores have their own place. When a measure has no norm table, raw scores are the honest metric, and the Standards require that when raw scores are meant to be interpreted directly, their meanings and limits be described and justified as carefully as for scale scores (Standard 5.4). Never rename a raw count as a standard score, and never report both under a single heading in a table.
6 How was the score derived: full battery, substitution, proration or short form?
Report whether the composite came from the complete set of core subtests, from a substituted subtest, from a prorated sum, or from a short form, because each route adds measurement error that a full battery does not. The WAIS-IV manuals show the mechanics for one battery, and the rules differ by instrument, so read the manual of the edition you used. The technical manual allows a supplemental subtest to replace a core subtest when the core score is invalid because of administration errors, recent exposure to items, physical limitations or sensory deficits, or response sets, and it notes that some supplemental subtests are available only for ages 16:0 to 69:11. It adds that when an allowable substitution has been made, the possible introduction of additional measurement error should be evaluated and noted in the report.
The Spanish-language edition of the administration manual goes further on proration. We read it in the Manual Moderno edition of 2014, and our summary is a paraphrase. It says the Full Scale composite and the General Ability Index may use no more than two substitutions, that substitution is always preferred over proration because the result then rests on actual performance, and that proration should be limited to cases where it cannot be avoided. It describes proration as a departure from standardized administration that may add measurement error, calls it particularly problematic when results are used for diagnosis, and instructs examiners to mark a prorated score on the record form. It allows a verbal or perceptual index to be prorated when two of its three core subtests are valid, does not allow the two-subtest working memory and processing speed indices to be prorated when one subtest is invalid, and permits a prorated Full Scale composite from nine or eight subtests.
For a methods section, three consequences follow.
Count the cases. State how many participants had a substituted or prorated composite and which subtest was involved. If the number is zero, say so.
Do not merge them silently. A prorated total is an estimate of the composite, not the composite. Report the main analysis with full scores only and run a sensitivity analysis with the derived scores included. This is our recommendation, built on the manual statement that these routes add error.
Label the score. Write "Full Scale IQ (prorated from nine subtests)" in the table, and keep the label in figure captions.
Short forms are the broader version of the same problem. The Standards observe that an assessment shortened by decreasing the number of items or tasks is likely to have lower reliability, and that a reliability evaluation applies to a particular assessment procedure and may change if the procedure changes substantially. They also say that when a test is built from a subset of the items of an existing test, evidence should be provided that scale scores, cut scores or norms are not distorted for the different versions (Standard 5.19). A short form that reports an estimated composite from a subset of subtests is therefore a different procedure with its own reliability, and its label should say so, as in "estimated Full Scale IQ from a four-subtest short form (publisher conversion table)". The page on what a short battery costs you explains the trade in plain terms.
The same logic applies to a battery that offers several forms of different length. If a study uses the shorter form, the paper should name it, give its number of subtests, and report the reliability and the interval for that form rather than for the full one.
7 What must the paper say about reliability and the confidence interval?
Report reliability for the scores you analyzed in your own sample, and report the interval around an individual score separately, because a manual coefficient describes the publisher sample and not yours. JARS-Quant asks authors to estimate and report reliability coefficients for the scores analyzed, meaning the researcher sample, where possible. It lists test-retest coefficients for longitudinal designs, internal consistency coefficients for composite scales, and interrater reliability for subjectively scored measures, which covers verbal tasks scored by raters. It also asks that, when reliability or validity coefficients come from other samples, such as those in a test manual, the basic demographic characteristics of those samples be reported.
The Standards say the same from the measurement side. Estimates of reliability and precision should be reported for each score to be interpreted (Standard 2.3), and the sampling procedures used to select test takers for reliability analyses, with descriptive statistics on those samples, should be reported (Standard 2.19). Two points in Chapter 2 matter for research samples. A comment notes that reliability and generalizability coefficients can differ substantially when groups differ in variance on the construct, while the standard error of measurement is less affected, so a restricted sample such as a student volunteer pool may show a lower coefficient than the manual even when the instrument is unchanged. And when significant variations in administration are permitted, separate reliability analyses should be provided for scores produced under each major variation if samples allow (Standard 2.10). An unsupervised online administration is such a variation relative to a standardization done with examiners, which is why the manual coefficient cannot simply be borrowed. The page on reliability and validity explains the distinction between the two ideas in plainer terms.
Two intervals get confused in papers, and a methods section should keep them apart.
The interval around one person's score. It comes from the standard error of measurement. The WAIS-IV technical manual gives a worked example: a Full Scale IQ of 106 with a standard error of measurement of 2.12 yields a 95 percent interval of 106 plus or minus 1.96 times 2.12, which the manual rounds to 102 to 110. The same manual states that its composite intervals were computed from overall average reliabilities and standard errors, because the age-specific values were very similar. In a paper, give individual intervals only where individual scores are discussed, such as case descriptions, and say which publisher table or formula produced them.
The interval around a group mean. It comes from the sample standard deviation and sample size, and it is a statement about sampling error, not measurement error. JARS-Quant asks for confidence intervals with effect sizes in the findings, and those are this second kind.
Do not substitute one for the other. A tight interval around a group mean says nothing about how precisely any participant was measured, and a wide individual interval does not make a group mean imprecise.
When item-level data exist, compute internal consistency in your sample and report the coefficient and its type. When only subtest scores are available, the options are narrower: report test-retest if you tested twice, and otherwise cite the manual value together with the demographics of the manual sample, as JARS-Quant requires, and say plainly that the value was not estimated in your data. That plain sentence is better than an unlabeled number, and it is our recommendation rather than a quoted rule.
8 How should administration mode, proctoring and exclusions be written down?
Name the administration mode with a standard vocabulary, describe the conditions, and state the exclusion rules and counts, because unsupervised administration and post hoc exclusions both change what a score means. The International Test Commission guidelines on computer-based and internet delivered testing, published in the International Journal of Testing in 2006 and read here in the ITC PDF, define four modes. Open mode has no direct human supervision and no way to authenticate the test taker. Controlled mode has no direct supervision but limits the test to known test takers, usually through a login. Supervised or proctored mode has a level of direct human supervision, and managed mode has a high level of supervision and control, typically in dedicated testing centers. Using these four words gives a reader a compact and shared description. Write "controlled mode, unsupervised, on participants own devices" instead of "online".
The Standards explain why the label matters. The chapter on psychological testing and assessment says the interpreter of results should be informed if a test was unproctored or administered under nonstandardized procedures. The comment to Standard 3.4 notes that unproctored administrations, where standardization cannot be ensured, could provide an advantage to some test takers, and that results should be interpreted with caution. Standard 1.10 asks that the conditions under which data were collected be described in enough detail for users to judge their relevance to local conditions, and a comment names the mode (unproctored online versus proctored on-site) among the features that can differ. Standard 9.9 asks that a user who alters the mode of administration have a sound rationale and, where possible, empirical evidence that reliability and validity are not compromised. The ITC guidelines add, in their section on computer-generated interpretive reports, that open or controlled testing may happen under nonstandardized or unproctored conditions, whereas score interpretations are often based on proctored, standardized administration.
A mode paragraph should carry the following.
Mode and supervision. The ITC term, and whether anyone observed the session.
Device and environment rules. What participants could use, whether timed tasks were enforced by software, and whether breaks were allowed.
Language and instructions. The language of delivery and any change to standard instructions.
Setting and dates. JARS-Quant asks for the settings and locations where data were collected and the dates of collection.
Disruptions. Standard 6.3 says changes or disruptions to standardized procedures should be documented, and its comment adds that records should be kept so that researchers can use only the standardized records if they wish.
Exclusions need the same discipline. JARS-Quant separates two things that papers often blend: inclusion and exclusion criteria, which define who is eligible, and data diagnostics, which include the criteria for excluding participants after data collection, the handling of missing data and the treatment of outliers. State both kinds, say which were set before the data were seen, and give the count lost at each rule. A flow of numbers (recruited, completed, excluded by rule, analyzed) lets a reader judge the sample behind the IQ mean. The page on online testing data quality covers how to build the rules for careless responding, and the page on time limits explains why timing rules change what a timed task measures.
Missing data deserve one sentence of their own. If a participant skipped a subtest, the derived score is a substituted or prorated one, and the earlier section on derivation applies. If a participant never started, they are an attrition case, and JARS-Quant asks that it be counted.
9 Why do old norms inflate scores, and what should a paper say about the Flynn effect?
Norms age because test performance in the population has drifted upward, so a score computed from old norms sits higher than the same performance would against new ones. The best single estimate comes from the meta-analysis by Trahan and colleagues (2014) in Psychological Bulletin. Across 285 studies (N = 14,031) since 1951 with administrations of two intelligence tests with different normative bases, the meta-analytic mean was 2.31 standard score points per decade, with a 95 percent interval of 1.99 to 2.64. For 53 comparisons involving modern Stanford-Binet and Wechsler tests, since 1972, it was 2.93 points per decade (interval 2.3 to 3.5), which the authors describe as comparable to earlier estimates of about 3 and not consistent with a diminishing effect. They define the Flynn effect as the observed rise in IQ scores over time, which results in norms obsolescence. Flynn's 1984 paper on the mean IQ of Americans from 1932 to 1978 is the classic early documentation. The page on the Flynn effect tells the longer story.
The size is not constant across abilities. Pietschnig and Voracek (2015) analyzed 271 independent samples, almost 4 million participants and 31 countries from 1909 to 2013, and estimated annual gains of 0.41 IQ points for fluid tests, 0.30 for spatial, 0.28 for full scale and 0.21 for crystallized tests. Those annual figures are 4.1, 3.0, 2.8 and 2.1 points per decade by our arithmetic. They also found gains stronger for adults than children and decreasing in more recent decades. And the direction can reverse: Bratsberg and Rogeberg (2018) used Norwegian conscription data for male birth cohorts from 1962 to 1991 and found that the rise, turning point and decline of cohort scores were all recoverable from within-family variation, which they read as environmental causes.
For a thesis, the arithmetic is simple and the caveats matter. WAIS-IV norms were collected from March 2007 to April 2008, and a 2026 administration is about 1.8 decades later. At 2.31 points per decade that is about 4 points, at 2.93 about 5 points, and at the rule of thumb of 3 points per decade quoted by Farmer and Floyd, who illustrate it with a test normed in 2006 compared with one normed in 2016, about 5 points. All three are our arithmetic, and none is a correction factor for a particular sample. The rate differs by domain, it has slowed in some periods, and it has reversed in at least one country, so a fixed adjustment is a convention, not a measurement.
What a paper should say follows from that.
State the norm year and the years elapsed. "Norms collected in 2007 and 2008; data collected in 2025" does the job.
Say what kind of claim depends on the norms. Absolute statements such as "our sample was above the population mean" depend on the norm vintage, while differences between groups scored against the same table do not depend on it in a simple way.
Do not adjust silently. If you subtract a Flynn correction, name the rate and its source, and report the unadjusted result beside the adjusted one.
Group studies by norm vintage when comparing across papers. Two studies reporting mean IQs from norms 20 years apart are not on the same footing, even when the instrument name is the same.
The Standards place the duty where this page does. Publishers should renorm with sufficient frequency (Standard 5.11), and users are responsible for not applying out of date norms. A methods section that gives the norm year lets the reader apply that judgment.
10 Which reporting errors make an IQ score unreadable?
Most reporting errors come from dropping one of the eight facts, and each has a one line repair. The table collects the errors that follow from the sources above. It is our synthesis, not a survey of reviewer behavior.
Error
Why it misleads
What to write instead
A prorated or short-form total reported as "Full Scale IQ"
Proration and shortening add measurement error and depart from the standardized procedure
Label the score as prorated or estimated, name the subtests used, count the cases
Averaged percentile ranks
Percentiles are not equal-interval, so their mean belongs to no scale
Average standard scores, then convert, or report the median percentile
An index or subtest score called "IQ"
The score covers fewer tasks and has its own reliability
Use the publisher label and reserve IQ for the global composite
State collection years; analyze editions separately or note nonequivalence
Manual reliability presented as sample reliability
The coefficient describes another sample, and variance affects it
Estimate in your data, or label it as a manual value with its sample
Individual interval used for a group mean, or none given
Measurement error and sampling error are different quantities
Give each interval its own name and source
A raw or within-sample score called IQ
No norm table converts it to the population metric
Call it a raw score or a z score within the sample
Mode left as "online"
Unproctored and proctored administrations are not interchangeable
Use the four ITC mode terms and describe conditions
Two of these deserve a further note. The first is the pooling of editions, such as scores from the WAIS-IV and the WAIS-5 in one analysis. The Standards say converted scores for two versions may not be strictly equivalent even after statistical linking (Standard 5.20), so the paper should either analyze them separately or defend the pooling with evidence. The second is treating a within-sample z score as IQ. A measure such as the ICAR can be useful for research, and its scores can be analyzed well, but unless a norm table exists, the honest label is a raw or z score.
A good instrument paragraph is short, fills all eight facts, and reads the same whether the measure is a published battery, an item pool or a short form. The three templates below use brackets for what you must fill in. They are our drafting aids, built from the requirements cited above, and a journal's own instructions override them.
Case 1: a normed battery taken online without supervision. "Cognitive ability was measured with [instrument, edition, form] ([publisher]). Participants completed it between [date] and [date] in [open or controlled] mode, without supervision, on [devices], with software-enforced time limits. Raw scores were converted to scaled scores (mean 10, SD 3) and composite standard scores (mean 100, SD 15) using the publisher adult norm table, collected in [years]. We analyzed the [Full Scale IQ or named index]. Reliability in this sample was [coefficient and type]. Individual 95 percent intervals use the publisher standard errors. We excluded [n] participants by [rules set before analysis], leaving [n]. No composite was prorated or substituted."
Case 2: a public-domain item pool without norms. "Reasoning was measured with [number] items from the International Cognitive Ability Resource (Condon and Revelle, 2014), namely [item types and counts], administered in [mode] between [dates]. No population norm table was applied. Scores are the number of items answered correctly and are not IQ scores; for analysis they were standardized within the sample. Internal consistency in this sample was [coefficient]. Exclusions were [rules and counts]." The ICAR paper describes a public-domain measure, and our page on the ICAR states that no peer reviewed publication or project page converts an ICAR raw score to an IQ metric or percentile, so check the project documentation for any later norm table before you write the sentence.
Case 3: a short form administered by an examiner. "[Short form] ([n] subtests) was administered individually by [qualified examiner] in [setting]. An estimated Full Scale IQ was derived with the publisher conversion table. It is an estimate from the short form and not the full battery composite. Reliability for the short form is the publisher value from a sample of [demographics] and was not estimated in our data. Norms were collected in [years]. Exclusions were [rules and counts]."
The table shows what each kind of instrument gives you to cite, and what you must supply yourself.
Instrument type
Norm table and date
Reliability and interval
What you must still report
Published full battery (WAIS-IV as the worked case)
Yes, with collection dates (March 2007 to April 2008)
That scores are not IQ, item types and counts, sample reliability
Short form of a battery
Publisher conversion, lower precision
Short-form value, if the publisher gives one
That the score is an estimate, subtests used
Normed online battery (ACIS)
Documented in the technical manual
The report shows a 95 percent interval
Form used, mode, sample reliability, norm documentation citation
12 Where ACIS sits, and what a buyer should do
ACIS is an online, unsupervised, normed battery for adults, so a study that uses it must report the mode plainly and must not present its scores as clinical or as WAIS scores. ACIS has 20 subtests in six CHC domains, reported as a Full Scale IQ and six indices (VCI, FRI, QRI, VSI, WMI, PSI), with scaled subtest scores from 1 to 19 and index and Full Scale scores on the standard scale with a mean of 100 and a standard deviation of 15. The report gives percentiles and a 95 percent confidence interval, and adult norms cover ages 16 to 90. It comes in three forms: Quick with 6 subtests in three domains (about 45 minutes), Optimized with 13 subtests in five domains (about 110 minutes), and Full Scale with all 20 subtests in six domains (about 175 minutes). Prices read on October 6, 2026 were 15, 30 and 50 dollars, as one time payments with no subscription, and prices can change.
Its limits are as important. It is online and unsupervised, it is not a clinical or diagnostic instrument, it is not for hiring, school accommodations or admission to high IQ societies, and it is available in English only. It is not a WAIS and does not produce WAIS scores. The page on professional versus online IQ tests and the page on the accuracy of online IQ tests discuss that gap in general.
For a researcher, the eight facts apply as follows.
Instrument and form. Name ACIS and the form used, with its subtest count.
Norms. Cite the technical manual for the norm documentation, and do not paraphrase norm details from this page, which holds none. Add the year of your own data collection.
Score level and scale. Say whether you analyzed the Full Scale IQ, a named index or subtest scaled scores, and keep percentiles for description.
Derivation. Say which form each participant took, since a Quick form does not contain all 20 subtests.
Reliability. Estimate it in your data. This page quotes none for ACIS, and the manual describes its own sample.
Mode. Choose the ITC term from how your study gives participants access, and report device and environment rules. The page on ACIS for research explains the research workflow.
Exclusions. Declare rules and counts, as above.
Comparability. If you also use a supervised battery, do not pool the scores without saying so.
The professional framework is the one this page has followed. The Standards for Educational and Psychological Testing (AERA, APA and NCME, 2014) place on test users the duties to study the developer materials (Standard 9.2), to document the rationale when a test is used for a purpose with little validity evidence (Standard 9.4), to justify a change in mode with evidence where possible (Standard 9.9), and to avoid out of date norms (Standard 5.11). The APA's JARS-Quant table supplies the manuscript checklist that carries those duties into a paper. A researcher who follows both can describe an ACIS score as precisely as any other, and a reviewer can then decide what it supports.
Every entry below was opened or checked before it was cited. Crossref confirmed each DOI record, the Standards, JARS-Quant table and ITC guidelines were read in full text copies, and the manuals and the Guilford chapter were read in local copies.
AERA, APA and NCME (2014). Standards for Educational and Psychological Testing. Washington, DC: American Educational Research Association. Open access edition, read for Chapters 1, 2, 3, 5, 6, 7, 9 and 10.
Appelbaum, M., Cooper, H., Kline, R. B., Mayo-Wilson, E., Nezu, A. M., and Rao, S. M. (2018). Journal article reporting standards for quantitative research in psychology: The APA Publications and Communications Board task force report. American Psychologist, 73(1), 3-25. doi.org/10.1037/amp0000191
American Psychological Association (2020). JARS-Quant Table 1: Information recommended for inclusion in manuscripts that report new data collections regardless of research design.PDF copy hosted by SAGE Publications.
International Test Commission (2006). International guidelines on computer-based and internet-delivered testing. International Journal of Testing, 6(2), 143-171. doi.org/10.1207/s15327574ijt0602_4
Trahan, L. H., Stuebing, K. K., Fletcher, J. M., and Hiscock, M. (2014). The Flynn effect: A meta-analysis. Psychological Bulletin, 140(5), 1332-1360. doi.org/10.1037/a0037173
Pietschnig, J., and Voracek, M. (2015). One century of global IQ gains: A formal meta-analysis of the Flynn effect (1909 to 2013). Perspectives on Psychological Science, 10(3), 282-306. doi.org/10.1177/1745691615577701
Bratsberg, B., and Rogeberg, O. (2018). Flynn effect and its reversal are both environmentally caused. Proceedings of the National Academy of Sciences, 115(26), 6674-6678. doi.org/10.1073/pnas.1718793115
Flynn, J. R. (1984). The mean IQ of Americans: Massive gains 1932 to 1978. Psychological Bulletin, 95(1), 29-51. doi.org/10.1037/0033-2909.95.1.29
Condon, D. M., and Revelle, W. (2014). The international cognitive ability resource: Development and initial validation of a public-domain measure. Intelligence, 43, 52-64. doi.org/10.1016/j.intell.2014.01.004
Wechsler, D. (2008). WAIS-IV technical and interpretive manual. San Antonio, TX: Pearson. Publisher product page. Local copy read for norm collection dates, composite structure, substitution and the confidence interval example.
Wechsler, D. (2014). WAIS-IV: Escala Wechsler de Inteligencia para Adultos, manual de aplicacion (Spanish-language edition, adapted from the Pearson 2008 original). Mexico City: Editorial El Manual Moderno. The English manuals are listed on the same Pearson product page. Local copy read for the proration and substitution rules.
Farmer, R. L., and Floyd, R. G. (2018). Use of intelligence tests in the identification of children and adolescents with intellectual disability. In D. P. Flanagan and E. M. McDonough (Eds.), Contemporary intellectual assessment: Theories, tests, and issues (4th ed., Chapter 23, from p. 643). New York: Guilford Press. Publisher page. Local copy read for the norm recency passage.
14 Frequently Asked Questions
How do you report an IQ score in a research paper?
State the score with its instrument, edition and form, the norm table and the year its data were collected, the score level and scale, how it was derived, the reliability in your sample, the administration mode and the exclusions. A bare number cannot be judged or compared by a reader.
How do you report Full Scale IQ in APA style?
Write the score label with the mean, standard deviation and sample size of the standard scores, then give the edition and norm year in the methods. APA JARS-Quant adds reliability for the scores analyzed, so report that too, and keep the label exact, for example Full Scale IQ.
What should the methods section say about an IQ test?
It should name the publisher, test, edition and form, the norms, the mode and conditions, the date range, the scores analyzed, how they were derived, reliability in your sample, and exclusion rules with counts. Describe task formats if useful, but never print items.
Do you need to report the norm year of an IQ test?
Yes. Norms describe the reference group at the time of testing, and the Standards ask that norming reports give the dates of testing. A reader needs the year to judge how far the norms may have drifted from the population you tested.
Should you report the IQ test version or edition?
Yes. Editions differ in subtests, composites, norms and sometimes names, and the Standards say converted scores across versions may not be strictly equivalent. Writing only the brand name hides which measurement produced the number and blocks fair comparison with other studies.
How do you report IQ scores in a thesis or dissertation?
Follow the same checklist as a journal article, then add detail the thesis format allows: the full norm table citation, the scoring software, a flow of exclusions, and a table of reliability by score level. Check your department guidelines, which can add requirements beyond these.
Do you report IQ as a mean and standard deviation?
For group analyses, yes: report the mean, standard deviation and sample size of standard scores, plus the confidence interval of the mean. Check that the sample standard deviation is not far below the scale value of 15, since a narrow spread can signal range restriction.
What does JARS-Quant say about the reliability of measures?
It asks authors to estimate and report reliability coefficients for the scores analyzed in their own sample where possible, covering test-retest, internal consistency and interrater types as relevant. If values come from a manual, it asks for the demographics of that sample.
Is it acceptable to report a prorated Full Scale IQ?
It can be acceptable if labeled. Publisher rules allow proration in limited cases, but they treat it as a departure from standard procedure that adds error. Label it as prorated, name the subtests used, count the cases, and show results without them.
Why can you not average percentile ranks?
Percentile ranks are not equal-interval, so equal changes in standard score cover very different shares of people. In a normal distribution, averaging the 50th and 97.7th percentiles gives 73.9, which matches a standard score near 110, not the 115 you get by averaging scores.
How many IQ points does the Flynn effect add per decade?
A meta-analysis of 285 studies found 2.31 standard score points per decade, and 2.93 for modern Stanford-Binet and Wechsler comparisons. Other work finds different rates by ability type and period, so treat any single number as an average rather than a fixed constant.
How do you report a confidence interval for an IQ score?
Give the level and limits, for example 95 percent CI 107 to 117, and say it comes from the standard error of measurement of the publisher or your own reliability. Keep it separate from the confidence interval of a group mean, which reflects sampling error.
What are the four ITC administration modes?
The International Test Commission names open mode with no supervision and no identity check, controlled mode with no supervision but known test takers, supervised or proctored mode with direct human supervision, and managed mode with a high level of control, usually in testing centers.
Does the 2014 Standards document require saying a test was unproctored?
It says the interpreter of results should be informed if a test was unproctored or run under nonstandardized procedures. It also asks that data collection conditions be described well enough for users to judge their relevance. In a paper, that means naming the mode explicitly.
Can you cite a manual reliability instead of computing your own?
You can cite it, but JARS-Quant prefers reliability for the scores analyzed, and it asks for the manual sample demographics when you borrow values. A manual coefficient describes another sample and administration, so label it clearly as such and note the differences.
How should you report IQ from a test with no norms, such as the ICAR?
Report raw scores or within-sample z scores and say they are not IQ scores, because no norm table converts them to the population metric. Give item types and counts, the mode, and reliability in your sample, and check the project documentation for any later norms.
Should you correct IQ scores for the Flynn effect?
Only if your question needs absolute levels, and then transparently. Name the rate and its source, and report unadjusted results beside adjusted ones. Group comparisons scored against the same norms are less affected, and the rate varies by domain and period.
Can you pool WAIS-IV and WAIS-5 scores in one analysis?
Only with a stated rationale. Standards guidance says converted scores for two versions may not be strictly equivalent even after linking, so analyze editions separately, include edition as a factor, or cite evidence supporting equivalence. Silent pooling hides a difference in norms and tasks.
How should exclusions be described in an IQ study?
Give the rule, whether it was set before seeing data, and the number removed at each step, from recruited to analyzed. JARS-Quant separates eligibility criteria from post-collection exclusion criteria and also asks how missing data were handled, so a complete report states both kinds of rule.
Can an online unsupervised IQ measure be used in published research?
Nothing in these sources forbids it, but the paper must say the administration was unsupervised and discuss what that changes. The Standards warn that unproctored conditions cannot ensure standardization, so results call for caution and for reliability estimated in your own sample.
Can I report ACIS scores in a research paper?
You can report them with the form used, the mode and your own reliability, citing the technical manual for norm documentation. ACIS is online, unsupervised and not a clinical or diagnostic instrument, so the paper should not present its scores as clinical or as WAIS scores.
Take the assessment
You get a profile, not a number
ACIS measures six CHC domains across 20 subtests and reports each one with its own normed score and confidence interval, so you can see where you are strong and where you are not.