Fairness and Validity

Are IQ Tests Biased? Six Questions, Not One

An IQ test can be biased in item behavior, construct representation, prediction or use. It can also be administered unfairly without being statistically biased, and equal average scores are not the definition of fairness. A defensible answer must identify which claim is being tested.

IQ test bias framework showing DIF, culture, validity, access and consequences
Bias is a technical claim about score meaning and use, not simply the observation that groups have different averages.

0 The Short Answer

Yes, an IQ test can be biased, and no, that claim does not mean what most people assume it means. In psychometrics, bias is a property of the instrument, not a verdict on the people who take it. A test is biased when it measures something different, or predicts an outcome differently, for two people who have the identical underlying level of the ability it claims to assess. It is not biased simply because average scores differ between groups. Those are two separate questions, and conflating them is the single most common error in public arguments about testing.

The reason this distinction matters is that the first question, whether the instrument itself is doing its job evenly, has a real answer that statisticians can compute. Item by item, a test publisher can check whether a question behaves the same way for test takers who are matched on ability but differ in group membership. A test can then be revised, an item can be dropped, and the fix can be verified on the next sample. The second question, why average scores differ between groups in the first place, is a much larger and more contested question about causes, and it is not the subject of this page.

This page stays on the first question and walks through the actual toolkit: differential item functioning, measurement invariance, predictive bias, content and language review, and the practical limits that online administration adds on top of all of it. Every method below is decades old, widely published, and used routinely by test publishers who take fairness seriously.

Where ACIS stands ACIS is not presented as culture free, and no test honestly can be. Scores are interpreted against ACIS's own norm sample, in the format ACIS actually uses, and with the limits of unsupervised online administration factored in. That is a narrower and more honest claim than "culture free," and it is the one this page defends.

1 What "Bias" Means in Psychometrics, Which Is Not What You Think

In everyday speech, calling a test biased usually means "this test produced a result I find unfair," often inferred straight from a gap in average scores between groups. Psychometrics uses the word much more narrowly, and the difference is worth sitting with because it changes what evidence would actually settle the argument.

A test item or a full test is biased, in the technical sense, when test takers of equal standing on the trait being measured have unequal chances of answering correctly, or when the same score corresponds to different real-world outcomes, because of group membership rather than the ability the test intends to capture. The key phrase is "equal standing on the trait." Bias analysis always starts by matching people on ability first, then checks whether anything besides ability is still predicting the outcome. A raw gap between group averages, with no such matching, tells you almost nothing about bias, because unequal average ability is exactly what an unbiased test would also produce if the groups genuinely differ in the trait it measures.

This is the concept researchers call construct-irrelevant variance: differences in test performance driven by something other than the construct the test is supposed to measure, for example unfamiliar vocabulary in a reasoning item, a figure drawn in a way that reads differently across cultures, or a word problem set in a context one group encounters more often than another. Construct-irrelevant variance is the actual target of every method described on this page. None of them can adjudicate whether observed group differences in average scores reflect real differences in the underlying trait, environmental factors, measurement error, or some mixture, and none of them try to. They answer a narrower, checkable question: is this specific item, or this specific score, doing something other than measuring the construct for some test takers?

The Standards for Educational and Psychological Testing, published jointly by the American Educational Research Association, the American Psychological Association, and the National Council on Measurement in Education in 2014, devote a full chapter to fairness and define it in almost exactly these terms: a test is fair to the degree that it measures the same construct, with comparable reliability, for all examinees in the intended population, and yields scores whose interpretation is equally valid across those examinees. That is the definition the rest of this page works from.

2 Differential Item Functioning: How a Single Question Gets Flagged

Differential item functioning, almost universally shortened to DIF, is the workhorse method for catching biased items before a test ever ships. The logic is simple to state and rigorous to run. Take two groups of test takers who score the same on the total test, or on some other proficiency measure external to the item being checked. If the item is behaving properly, both groups should have roughly the same probability of answering it correctly, because they have already been matched on the ability the item is supposed to measure. If one group is still doing noticeably worse on that specific item despite equal overall ability, the item is flagged for DIF.

The oldest and still most widely used detection method is the Mantel-Haenszel procedure, adapted for testing by statisticians Paul Holland and Dorothy Thayer in a 1988 Educational Testing Service research report. The technique was originally built for medical epidemiology, comparing an exposure's effect across matched strata, and Holland and Thayer showed it could be repurposed to compare an item's difficulty across ability-matched groups in a 2x2xK contingency table. It remains popular because it needs no assumption about the shape of the underlying item response curve and performs well even with moderate sample sizes.

Item response theory offers a second family of methods, comparing the full item characteristic curve, not just a single pass rate, between groups. A likelihood ratio test can check whether an item's difficulty and discrimination parameters differ across groups once ability is held constant, and logistic regression approaches let researchers test for both uniform DIF, where one group is disadvantaged across the whole ability range, and nonuniform DIF, where the disadvantage appears only at certain ability levels and can even reverse direction elsewhere on the scale. The Mantel-Haenszel procedure is sensitive mainly to the uniform kind, which is one reason serious test development programs run more than one method rather than relying on a single flag.

Mantel-Haenszel

Nonparametric, contingency-table method comparing pass rates across ability-matched groups. Fast, well established, best at uniform DIF.

IRT likelihood ratio

Compares full item response curves between groups, catching both uniform and nonuniform DIF at the cost of needing larger samples.

Logistic regression

Models correct response as a function of ability, group, and their interaction, separating uniform from nonuniform DIF in a single test.

None of these methods explains why an item shows DIF. A flagged item might contain a word one group encounters less often, a picture drawn in an unfamiliar style, or a scenario tied to a specific cultural context. Statistics identify the symptom. A content review committee, usually a mix of psychometricians and subject matter reviewers from varied backgrounds, has to diagnose the cause and decide whether the item can be fixed or should be dropped.

3 How Test Publishers Actually Use DIF in Development

DIF analysis is not a one-time audit run after a test is finished. In a well-run development program it is built into the pipeline from the first field trial onward, and it shapes which items even survive to the published version.

The typical sequence starts with a large item pool, often two or three times the number of questions the finished test will need, administered to a tryout sample that is deliberately built to include enough people from each comparison group to make DIF statistics stable. Every item in the pool gets a DIF flag: green for no detectable difference, amber for a small effect worth a second look, red for an effect large enough to demand action. The classification thresholds themselves are published and standardized, most commonly the effect-size bands originally proposed for Educational Testing Service programs, so that a flagged item is not a judgment call made on the spot.

Items that flag red are not simply deleted by an algorithm. They go to a bias review committee, a group intentionally assembled to include reviewers from different backgrounds, who read the item alongside the statistics and try to identify the likely mechanism: an idiom, a specific cultural reference, an image that reads ambiguously, a context more familiar to some test takers than others. Some flagged items survive this review because the committee concludes the statistical flag reflects a genuine and defensible difficulty gradient rather than construct-irrelevant content. Others are rewritten and sent back through another tryout round. The remainder are dropped and never appear in a scored form of the test.

This is also why publishers periodically rebuild norms and re-run item analyses rather than treating a test as finished forever. Language drifts, cultural references age, and a word or image that was neutral at launch can pick up new associations a decade later. A DIF flag that would not have appeared in an original field trial can appear in a later cohort, which is one of the practical reasons psychometric batteries get revised editions on a schedule rather than staying static indefinitely.

4 Measurement Invariance: Proving the Test Measures the Same Thing

DIF analysis works item by item. Measurement invariance testing asks the same fairness question about the test as a whole, and specifically about the underlying structure the items are supposed to share. If a test claims to measure, say, fluid reasoning through a set of matrix items, invariance testing checks whether that entire set of items relates to the fluid reasoning factor the same way across groups, not just whether any single item is flagged.

The standard framework, laid out clearly in a widely cited 2016 methodological review by Diane Putnick and Marc Bornstein in Developmental Review, breaks the question into ordered steps, each one a stricter test than the one before. Configural invariance asks whether the same basic factor structure fits every group: the same number of underlying factors, with each item loading on the factor it is supposed to. Metric invariance goes further and asks whether the strength of that relationship, the factor loading, is the same across groups, which matters because it determines whether a one-point change in the trait produces the same change in observed score everywhere. Scalar invariance is the strictest practical level, requiring that item intercepts also match across groups, which is the condition that has to hold before comparing mean scores across groups means anything at all.

Configural

Same factor structure across groups: the same items load on the same underlying trait everywhere.

Metric

Same factor loadings across groups: a one-unit shift in the trait moves the observed score by the same amount everywhere.

Scalar

Same item intercepts across groups: the precondition for comparing mean scores across groups in a way that is statistically defensible.

Each step is tested with multi-group confirmatory factor analysis, fitting the model separately for each group and then comparing fit as constraints are added. If adding the metric constraint makes the model fit substantially worse, metric invariance fails and the test is, in a technical and checkable sense, not measuring the trait in a directly comparable way across those groups, whatever the raw scores suggest.

5 Why Invariance Testing Is a High Bar, and What Happens When It Fails

It is worth being honest about how demanding full scalar invariance actually is, because it is common for large, well constructed tests to achieve configural and metric invariance cleanly while a handful of items fail the scalar step. When that happens, researchers do not automatically throw out the whole instrument. The standard response is a model called partial invariance, in which most items are held to the strict scalar standard and the small number that fail are allowed their own intercepts, provided enough of the item set still meets the full standard to anchor the comparison.

The practical stakes are highest exactly where invariance is weakest: comparing average scores across groups. If scalar invariance does not hold, a difference in mean scores between two groups can no longer be cleanly attributed to a difference in the underlying trait, because part of that difference could be coming from the item intercepts themselves rather than from the trait the test is trying to compare. This is precisely why psychometricians treat cross-group score comparisons as claims that require invariance evidence attached, not as something a raw score table can support on its own.

This also explains why serious critiques of specific instruments tend to be narrow and technical rather than sweeping. A finding that one subtest fails scalar invariance between two specific groups, on one specific version of a test, is a real and actionable finding: it points at particular items, on a particular instrument, for a particular comparison. It is not evidence that intelligence testing as a category is unworkable, any more than one poorly calibrated thermometer means thermometers cannot measure temperature. The fix in both cases is the same: identify the faulty component and correct or replace it.

6 Predictive Bias: A Different Question From Item Bias

Everything so far concerns internal or structural bias: does the test measure the same construct, in the same way, for everyone. Predictive bias asks a separate and equally important question: once the test produces a score, does that score predict an external outcome, such as academic performance or job performance, equally well for every group. Admissions tests are where that question gets asked most often, and the ACT and IQ comparison keeps the two evaluations apart: the ACT is judged against outcomes such as course performance and college success, while a cognitive battery is judged against age based norms.

The classic framework here comes from psychologist T. Anne Cleary, who proposed in a 1968 Journal of Educational Measurement paper what is still called the Cleary model. The idea is to plot the test score against the criterion it is meant to predict separately for each group and compare the resulting regression lines. If the lines are effectively the same, a given test score predicts the same criterion outcome regardless of group membership, and the test shows no predictive bias by this definition. If the lines differ, they can differ in two distinct ways worth separating clearly. A difference in slope means the test's predictive strength itself varies by group, the same score change corresponds to a different-sized change in the outcome. A difference in intercept means the test systematically over-predicts or under-predicts the criterion for one group by a roughly constant amount, even when the slope is identical.

These are two genuinely different problems with different implications. Slope bias means the test is a worse predictor for one group across the board. Intercept bias means the test is a consistently miscalibrated predictor in one direction for one group, which, depending on which direction it runs, can either work against a group or, less intuitively, work in its favor by predicting a better outcome than actually materializes. Neither problem is fixed by adjusting item content the way DIF findings are. Predictive bias is a property of the score's relationship to an outside criterion, not of the items themselves, so the fix usually involves recalibrating the prediction model rather than rewriting the test.

7 What the Predictive Bias Literature Actually Shows, Stated Carefully

This section stays deliberately narrow, because predictive bias research is exactly the place where public discussion tends to slide from a checkable statistical question into a much larger and unsettled argument about the causes of group differences. That larger argument is outside the scope of a page about testing methods, and nothing here should be read as taking a side in it.

What the methodology itself supports saying is this: predictive bias, in Cleary's sense, is a testable property of a specific test used with a specific criterion in a specific setting, and it does not have to be assumed either present or absent. It has to be measured with paired score and criterion data, using the regression comparison described above, for each pairing of test and outcome a publisher wants to defend. A test validated for predicting first-year college grades has not thereby been validated for predicting job performance in an unrelated field, and a finding of no predictive bias in one context is not a blanket finding about the instrument everywhere it might be used.

Two methodological cautions come up repeatedly in the literature on this topic. The first is that predictive bias analysis is only as good as the criterion measure itself; if the outcome being predicted, such as a supervisor rating or a course grade, is itself measured inconsistently across groups, the bias analysis inherits that problem no matter how carefully the test score side is handled. The second is that omitted variables can produce the appearance of slope or intercept differences that have nothing to do with the test, which is why credible predictive bias studies control for the obvious confounds before concluding the test itself is the source of a discrepancy. Readers evaluating any claim about predictive bias, in either direction, should ask what criterion was used, how it was measured, and what else was controlled before accepting the conclusion.

8 Content Bias and the Special Problem of Verbal Items

Not every source of unfairness shows up cleanly as a DIF statistic on a finished item. Content bias is the broader, upstream concern: whether the pool of situations, words, and images a test draws from gives systematic advantages to test takers with certain backgrounds, independent of the ability the test is trying to measure. Content review usually happens before an item ever reaches a tryout sample, through panels that read every item for cultural specificity, idiom, regional vocabulary, and assumed background knowledge. How unevenly that exposure is spread becomes obvious once the six item families a full battery samples are set out side by side, since verbal and quantitative questions presuppose a particular language and a particular schooling in a way that spatial, working memory and processing speed tasks largely do not.

Verbal items carry disproportionately more of this risk than nonverbal ones, for a structural reason rather than a coincidental one. A vocabulary or verbal comprehension item is, almost by definition, testing familiarity with a specific language as it is actually used, and language exposure varies with education, region, native language, and socioeconomic background in ways that have nothing to do with fluid reasoning or working memory. A word that is common in one dialect or one country's schooling can be rare or unknown in another, which is exactly the kind of construct-irrelevant variance DIF and content review exist to catch. This is also the practical argument for building some subtests around fluid reasoning tasks, such as visual pattern completion, that lean less on a specific vocabulary and more on reasoning with novel material presented nonverbally.

Reduced-language and nonverbal formats are a genuine partial answer to this problem, not a complete one. Matrix reasoning and figural analogy tasks depend far less on a specific vocabulary, which is why they show up in most modern batteries designed for use across linguistic groups. But nonverbal does not mean assumption-free. Figural tasks still assume familiarity with two-dimensional printed diagrams, a left-to-right or top-to-bottom scanning convention, and the general test-taking format itself, which is a smaller but real kind of prior exposure that is not evenly distributed either. Reducing language load shrinks one category of construct-irrelevant variance; it does not remove every category at once.

9 Why "Culture Free" Is a Promise No Test Can Fully Keep

The phrase "culture free" has a specific, somewhat cautionary history in psychometrics. Raymond Cattell built his Culture Fair Intelligence Test in the mid-twentieth century specifically to strip out verbal and educational content and rely entirely on abstract figural reasoning, aiming for a measure that would work identically regardless of a test taker's cultural background. It remains a genuinely useful instrument and an influential design template. It was never able to fully deliver on the name. Its modern successor label inherits the same tension: the APA defines a nonverbal test purely as one whose questions and answers are not communicated in words, which describes a format and promises nothing about culture.

The clearest statement of why comes from developmental psychologist Patricia Greenfield, whose 1997 paper in American Psychologist, titled "You Can't Take It With You: Why Ability Assessments Don't Cross Cultures," argued that even entirely nonverbal, entirely abstract test formats carry embedded assumptions about the testing situation itself: what a test is for, how one is expected to behave when told to solve a puzzle for a stranger, how much weight to give speed versus accuracy, and what kind of reasoning is considered valid evidence of intelligence in the first place. These are not vocabulary problems, so language-free formats do not solve them. They are assumptions baked into the format of standardized testing as a social practice, and every standardized cognitive test, whatever its content, inherits them.

This is why the field has largely moved from claiming tests can be culture free to the more modest and more accurate claim that a test can be culture reduced, meaning it minimizes specific, identifiable sources of unfair advantage that content review and DIF analysis can actually detect, while acknowledging that the format of standardized testing itself is not neutral in some absolute sense. It is a smaller claim, and it is the honest one. A test that reduces detectable sources of unfairness through documented methods is doing real, verifiable work. A test that claims to have eliminated culture from the measurement altogether is making a claim psychometrics does not have the tools to support.

10 The Limits of Online, Unproctored Administration

Everything discussed so far concerns the test itself: its items, its structure, its predictive relationships. Online administration introduces a separate category of fairness concern that sits entirely outside the test content, in the conditions under which the test is taken, and it applies to essentially every unproctored online cognitive test, this one included. Anyone who needs the controlled version does have somewhere to go, and the venues that supervise administration, a licensed psychologist's office, a university training clinic, a hospital neuropsychology department or a proctored Mensa session, differ mainly in whether a documented report comes out at the end.

A supervised, in-person test administration controls variables that a home computer cannot: a quiet room, a fixed time limit enforced by an observer, a standard device with a known screen size and input method, and a test taker who cannot pause to look something up. None of that is guaranteed online. The International Test Commission's guidelines on computer-based and internet-delivered testing, first published in 2005 and revised since, flag exactly this gap, recommending that publishers using unsupervised online administration disclose the difference explicitly and avoid treating online scores as strictly interchangeable with proctored ones, precisely because the environment itself becomes a source of measurement error that a norm sample collected under different conditions cannot fully absorb.

Device and interface familiarity is a related and underappreciated piece of this. A timed test that depends on quick mouse clicks, drag actions, or a specific input pattern will run differently for someone who uses that kind of interface daily than for someone encountering it for the first time, and that gap has nothing to do with fluid reasoning, working memory, or the trait the test is trying to rank. Processing speed subtests are particularly exposed to this problem, because they are, by design, measuring how fast someone can execute a repetitive task, which makes them sensitive to interface friction in a way untimed subtests are not.

Distraction and self-pacing round out the list. A test taker interrupted mid-item, or one who takes a test in a noisy environment, is being measured under conditions the norm sample almost certainly did not experience uniformly. None of this means online testing is worthless, and unsupervised online batteries can still produce useful, internally consistent scores. It does mean online scores carry an extra layer of uncertainty that a published norm table, on its own, does not communicate, and any platform administering tests this way owes its users a plain statement of that limitation rather than silence about it.

11 How Serious Test Developers Actually Build and Audit for Fairness

Put together, the methods above form a pipeline rather than a single checkbox, and the sequence matters. Content review happens first, screening items for obvious cultural specificity, idiom, or assumed background knowledge before any statistics are even collected. Field tryout on a large, deliberately diverse sample comes next, producing the data DIF analysis needs. DIF flags send suspect items to a bias review committee, which decides whether to revise, retain, or drop each one, and the revised item pool is what actually becomes the scored test, not the original draft. Skipping that pipeline entirely is what defines the unregulated end of the market: the untimed take-home puzzle sets marketed as high range tests have no representative norm sample behind them, which means no field tryout, no DIF data and no bias review ever took place.

After a test is assembled, measurement invariance testing checks whether the finished structure holds together across groups, at the configural, metric, and, ideally, scalar level, and any predictive claims the publisher intends to make get their own separate predictive bias study against the actual criterion being promised, whether that is academic performance, job performance, or something else. The Standards for Educational and Psychological Testing treat this whole sequence, not any single step in isolation, as the baseline expectation for a defensible, well documented instrument, and independent test reviews, such as those published in the Buros Mental Measurements Yearbook, routinely check whether a publisher's technical documentation actually reports this evidence or simply asserts fairness without it.

None of this guarantees a perfect instrument. It guarantees a checkable one. The practical difference between a rigorously developed test and an unaudited one is not that the rigorous test has no flagged items. It is that the rigorous test can show its work: which items were flagged, what happened to them, what the invariance testing found, and what the predictive validity evidence actually covers. A test that cannot produce that paper trail has not been shown to be unbiased. It has simply not been checked, which is a different and considerably weaker claim than the marketing usually implies.

12 Where ACIS Stands on Its Own Limits

ACIS is a self-administered online cognitive assessment, not a clinical or diagnostic instrument, and it is built around six broad ability domains from Cattell-Horn-Carroll theory: fluid reasoning, crystallized or verbal comprehension, quantitative reasoning, visual spatial processing, working memory, and processing speed. Reporting a profile across those domains, rather than a single number, follows the same logic covered in the sections above: a single score collapses information that a domain profile keeps, and a profile is what lets a test taker and, if relevant, an interpreter see where the measurement is stronger and where it is weaker.

Consistent with everything on this page, ACIS makes no claim to be culture free, because no standardized test can honestly make that claim. Scores are meant to be read against ACIS's own norm sample, in the specific item formats ACIS actually uses, and administered the way ACIS actually administers them: online, self-paced, and unsupervised, with all the limits described in the section above on administration conditions. A result should be understood as a measurement taken under those specific conditions, not as an absolute, context-free statement about a person's ability. That is a narrower claim than many commercial tests make, and it is the accurate one.

If the mechanics in this page made you want to see what a domain-level profile actually looks like rather than a single headline number, that is something you can check directly.

13 Frequently Asked Questions

Are IQ tests biased?

Individual items and instruments can show measurable bias, which is why publishers run item-level and structural checks during development. Whether a specific test is biased is an empirical question answered by that test's own technical documentation, not a yes-or-no fact about IQ testing as a category.

Does a score gap between two groups automatically prove a test is biased?

No. A raw average difference says nothing about bias on its own, because psychometric bias requires comparing people who are already matched on ability. An unbiased test can still show average group differences if the groups genuinely differ in the trait being measured, and separating those two possibilities is exactly what DIF and invariance testing exist to do.

What is differential item functioning in plain terms?

It is a statistical flag raised when two test takers with equal overall ability have unequal chances of answering one specific question correctly. The flag points at a single item, not the whole test, and tells developers exactly where to look.

Who invented the main method for detecting DIF?

The most widely used method, the Mantel-Haenszel procedure adapted for test items, was described by Paul Holland and Dorothy Thayer in a 1988 Educational Testing Service report, building on a statistical technique originally developed for medical research.

What happens to a test question once it gets flagged for DIF?

It goes to a bias review committee, which reads the item alongside the statistics to identify a likely cause. Depending on the finding, the item is retained as is, rewritten and retested, or dropped from the scored version entirely.

What is measurement invariance?

It is a set of statistical tests, run at configural, metric, and scalar levels, that check whether a test's underlying structure works the same way across different groups of test takers, which is a precondition for comparing scores or means across those groups meaningfully.

What does it mean if a test fails scalar invariance?

It means item intercepts differ across groups, so comparing raw mean scores between those groups is not statistically defensible without further adjustment. Researchers typically respond with a partial invariance model rather than discarding the test outright.

What is predictive bias, and how is it different from item bias?

Item bias concerns whether the test itself measures the same construct fairly. Predictive bias concerns whether a resulting score predicts an outside outcome, such as academic or job performance, equally well across groups. A test can pass one check and still need the other.

Who developed the standard framework for predictive bias?

Psychologist T. Anne Cleary proposed the regression-based model in a 1968 Journal of Educational Measurement paper, comparing regression lines that predict an outcome from test scores separately for each group.

What is the difference between slope bias and intercept bias?

Slope bias means the test predicts the outcome with different accuracy across groups. Intercept bias means the test consistently over-predicts or under-predicts the outcome for one group by a roughly constant amount, even when accuracy itself is equal.

Why are verbal items considered more vulnerable to bias than nonverbal items?

Vocabulary and verbal comprehension items depend on familiarity with a specific language as it is actually used day to day, and that exposure varies by education, region, and native language in ways unrelated to the reasoning ability the test intends to measure.

Are nonverbal or matrix reasoning tests free of bias?

No test format is free of every source of bias. Nonverbal formats reduce language-related content bias but still assume familiarity with printed diagrams, standard scanning conventions, and the general testing situation itself, which is not evenly distributed either.

Is any IQ test truly culture free?

No credible psychometrician claims this today. Even Raymond Cattell's Culture Fair Intelligence Test, built specifically to remove verbal and educational content, could not eliminate the embedded assumptions of standardized testing as a format, a point argued influentially by Patricia Greenfield in a 1997 American Psychologist paper.

What does "culture reduced" mean if not culture free?

It describes a test that minimizes specific, detectable sources of unfair advantage through content review and statistical testing, while acknowledging that standardized testing as a practice carries assumptions no format removes entirely.

Does taking an IQ test online instead of in person introduce bias?

It introduces measurement error from a different source: an uncontrolled environment, variable device familiarity, and the absence of a proctor enforcing standard conditions. The International Test Commission's guidelines on internet-delivered testing specifically flag this as a fairness issue publishers should disclose.

Why would a timed test be more affected by online administration than an untimed one?

Timed subtests, especially those measuring processing speed, depend on how quickly a test taker can execute the required action, which makes them sensitive to interface familiarity and device responsiveness in ways an untimed reasoning item is not.

How do test publishers actually check fairness before releasing a test?

Through a sequence: content review of every item before tryout, statistical DIF analysis on a diverse field sample, bias committee review of flagged items, measurement invariance testing on the finished structure, and separate predictive bias studies for any outcome the test claims to predict.

What professional standard governs fairness in test development?

The Standards for Educational and Psychological Testing, published jointly by the American Educational Research Association, the American Psychological Association, and the National Council on Measurement in Education, most recently revised in 2014, with an entire chapter devoted specifically to fairness in testing.

Is ACIS a culture free test?

No, and ACIS does not claim to be. Scores are interpreted against ACIS's own norm sample, its specific item formats, and the conditions of unsupervised online administration, which is a narrower and more accurate claim than culture free.

Does ACIS report a single IQ number or a profile across domains?

ACIS reports a profile across six broad ability domains from Cattell-Horn-Carroll theory, fluid reasoning, verbal comprehension, quantitative reasoning, visual spatial processing, working memory, and processing speed, rather than a single collapsed figure.

If bias can be detected and fixed, why does the topic stay controversial?

Because the checkable, statistical question covered on this page, whether an instrument measures and predicts consistently across groups, is often discussed alongside a much larger and unsettled question about why groups differ in average scores at all. The two questions get argued together even though only the first one has methods that produce a clear, checkable answer.

Sources Behind This Page

The claims on this page follow the published literature rather than our own assertions. These are the primary papers and reference bodies worth reading directly, with what each contributes.

  • Nisbett, R.E. et al. (2012). Intelligence: new findings and theoretical developments. American Psychologist, 67(2). The broad APA review of what moves measured intelligence and what does not.
  • American Psychological Association. The Standards for Educational and Psychological Testing, the joint AERA, APA and NCME framework that legitimate tests are built and evaluated against.
  • Plomin, R. & Deary, I.J. (2015). Genetics and intelligence differences: five special findings. Molecular Psychiatry. Open access. The standard modern review of what twin and DNA evidence does and does not show.
  • Gottfredson, L.S. (1997). Mainstream Science on Intelligence, the editorial signed by 52 researchers. Intelligence, 24(1), hosted by the University of Delaware. A consensus statement on what IQ tests measure.
  • Spearman, C. (1904). General intelligence, objectively determined and measured. American Journal of Psychology, full text at Classics in the History of Psychology. The paper where the g factor entered psychology.
  • Voncken, L., Albers, C.J. & Timmerman, M.E. (2019). Improving confidence intervals for normed test scores. Behavior Research Methods. Open access. Documents the mean 100, SD 15 metric and the uncertainty that norming from samples adds to any score.
  • Pearson Clinical Assessment Scientific Council (2023). Standardized Clinical Assessment for Practitioners: A Primer. How standard scores, percentile ranks and the standard error of measurement are meant to be read together.
  • Pearson (2008). WAIS-IV Score Report sample. What a real report contains: every composite paired with a percentile rank, a 95% confidence interval and a qualitative description, never a bare number.
  • Trahan, L.H., Stuebing, K.K., Hiscock, M.K. & Fletcher, J.M. (2014). The Flynn effect: A meta-analysis. Psychological Bulletin. Open access. Scores rose about 2.31 points per decade across 285 studies, which is why norms age and older scores overstate standing.
  • Pearson (2024). WAIS-5, Wechsler Adult Intelligence Scale, Fifth Edition. The current adult battery, covering ages 16:0 to 90:11 across five cognitive domains.
Take the assessment

You get a profile, not a number

ACIS measures six CHC domains across 20 subtests and reports each one with its own normed score and confidence interval, so you can see where you are strong and where you are not.

Free trial, no card required. Full report from $15.