Online cognitive testing data quality: who you recruit changes the data more than the task does
Online cognitive data is cheap to collect and uneven in quality, and the unevenness follows the platform, the device and the absence of a proctor. This page sets out what Peer, Douglas, Uittenhove, Popovic, Steger and others measured, how bots and careless responding fit in, which screening rules the evidence and the ITC and ATP guidelines support, and what each rule costs.
The photograph is illustrative; in the 2023 working memory study discussed below, participants used their own laptop or desktop computers, and 71.3 percent of the Prolific sample passed every quality criterion.
0 The short answer
Online cognitive data quality depends first on who is recruited and second on how the task is delivered, and exclusion rates swing widely with both the pool and the rule. In a 2023 working memory study, 34.7 percent of MTurk participants passed every quality criterion against 71.3 percent of Prolific participants, and in a 2025 face recognition study, stricter attention checks excluded about 62 percent of MTurk participants against about 22 percent on Prolific. Unproctored ability testing adds a separate layer: a meta-analysis of 49 studies found a pooled standardized difference of 0.20 in favor of unproctored scores. Decide the exclusion rules before you see the data, apply the same rules to every group, and report what each rule removed.
DisclosureACIS sells a competing paid assessment, so we have a commercial interest in how this page reads. Every figure below comes from the paper or guideline named beside it, and arithmetic we did ourselves is labeled as ours. This page describes no ACIS screening method, norm or statistic. The only ACIS facts are the product facts near the end, read on October 6, 2026.
34.7 vs 71.3
Percent of MTurk and Prolific participants whose working memory data passed every quality criterion in Uittenhove, Jeanneret and Vergauwe (2023).
62 vs 22
Approximate percent of MTurk and Prolific participants excluded by stricter attention checks in a face recognition study, Popovic and colleagues (2025).
0.20
Pooled standardized difference in favor of unproctored over proctored ability scores across 49 studies, with a 95 percent confidence interval of 0.10 to 0.31 (Steger and colleagues, 2020).
1 What Changes When Nobody Supervises a Cognitive Test?
When nobody watches, five separate things stop being controlled: who the participant is, how much effort they give, what device and room they use, what help they can reach, and whether they have seen the material before. A proctored session controls these by presence. An online session controls them, if at all, by design and by screening afterward, and each one leaves a different trace in the data. The 2022 Guidelines for Technology-Based Assessment, published by the International Test Commission (ITC) and the Association of Test Publishers (ATP) and read on October 6, 2026, name the same territory in their delivery and security chapters: environmental conditions, authentication of identity, monitoring, device differences, response time recording and the detection of cheating.
The ITC and ATP guidelines are explicit that supervision is a matter of stakes. Guideline 3.14 says testing should be conducted under appropriate environmental conditions and that outcomes should be interpreted in light of those conditions. Its comment adds that for "unproctored, self-administered testing in a low-stakes environment," guidance should be given to the test taker on the required conditions and procedures. Guideline 3.15 asks for guidance on the required level of authentication of identity, and the comment says the appropriate level depends on the nature and stakes of the exam. Guideline 3.16 asks that procedures be monitored so that security is maintained, "through in-person proctoring or supervision, or by using remote monitoring through cameras if the stakes of the assessment warrant." A study that pays participants a few dollars for a puzzle task sits at one end of that scale. A selection decision sits at the other.
The Standards for Educational and Psychological Testing (AERA, APA and NCME, 2014), which the American Psychological Association co-publishes with the other two bodies, say the same thing from the side of validity. In the chapter on fairness they warn that "unproctored test administrations where standardization cannot be ensured, could provide an advantage to some test takers over others," and that in such situations "test results should be interpreted with caution." In the chapter on psychological testing they say the interpreter of the results should be informed if the test was unproctored. In the chapter on validity, they list the mode of administration, with the example of unproctored online testing versus proctored on-site testing, among the conditions that may change whether validity evidence carries over to a new setting. A design that treats online data as interchangeable with laboratory data has skipped that step.
None of this means online data are poor by default. The third section shows that one well-run online pool was almost indistinguishable from web-tested university students on a working memory task. The point is narrower: quality is a property of a particular pool, task, device mix and time, so it has to be measured in your sample rather than assumed from the platform name. The companion page on choosing an intelligence measure for an online study (see also what IQ tests measure) covers the prior question of which instrument to use. This page covers what happens to the scores after you have chosen.
2 How Much Does the Recruitment Platform Change the Data?
In a working memory study that separated the two, the participant pool changed the data far more than the delivery mode, with Prolific close to web-tested students and MTurk far from both. Uittenhove, Jeanneret and Vergauwe (2023, Journal of Cognition) tested 751 participants on a working memory task built on the Brown-Peterson design: 196 from MTurk, 300 from Prolific, 215 web-tested university students and 40 laboratory-tested students. Participants used their own laptop or desktop computers. The MTurk sample was screened for an approval rating above 95 percent and at least 100 completed tasks; the Prolific sample was not pre-screened. The authors summarize the result in their title: who you test is more important than how you test.
The numbers behind that title are the clearest in the literature for a cognitive task. The share of participants whose data passed every quality criterion was 87.6 percent for laboratory students, 72.6 percent for web-tested students, 71.3 percent for Prolific and 34.7 percent for MTurk. The authors call the web-tested students' result a relative loss of 17.1 percent against monitored laboratory testing. Before the final criteria, anomalous response patterns removed 17.3 percent of the MTurk sample, 9.3 percent of Prolific, 8.4 percent of web-tested students and 7.5 percent of laboratory students. A benchmark effect the task is known to produce, the disruption of rehearsal by irrelevant speech, was absent for 42 percent of MTurk participants, 16.9 percent of Prolific, 16.2 percent of web-tested students and none of the laboratory students.
The same pattern appears across other measures and years, and the table collects the studies a researcher is most likely to meet. Each row reports only what the paper states; the last column is our summary of the limit the paper itself or its design imposes.
Study
Sources compared
Sample
Measure
What it found
Limit to keep in mind
Peer and colleagues, 2017, Journal of Experimental Social Psychology
MTurk, CrowdFlower, Prolific Academic
Two studies
Attention, naivety, honesty
Prolific data quality was higher than CrowdFlower's and comparable to MTurk's; both alternatives were more naive and less dishonest than MTurk
Describes the platforms as they were in 2017
Peer and colleagues, 2022, Behavior Research Methods
MTurk, CloudResearch, Prolific, Qualtrics, Dynata
Study 1 final 2,429; Study 2 final 1,461
Attention, comprehension, honesty, reliability
Prolific high on all measures unfiltered; CloudResearch and Prolific good with filters; MTurk poor even with filters
Surveys and honesty tasks, not a cognitive battery
Prolific and CloudResearch participants passed more checks and worked slowly enough to read the items
Survey experiment, not a cognitive battery
Uittenhove and colleagues, 2023, Journal of Cognition
MTurk, Prolific, web students, lab students
751
Working memory task with 70 trials
Passing all criteria: lab 87.6, web students 72.6, Prolific 71.3, MTurk 34.7 percent
One task; lab group only 40
Popovic and colleagues, 2025, Scientific Reports
MTurk, Prolific, laboratory students
Three face identity tests
GFMT2, CFMT+ and MFMT
MTurk about 10 percentage points lower; stricter attention checks excluded about 62 percent of MTurk and 22 percent of Prolific
Face identity ability, not general cognitive ability
Esch, Mylonopoulos and Theoharakis, 2025, Behavior Research Methods
MTurk, Prolific, Pollfish, Qualtrics
Two studies
Composite of attentiveness measures
Mobile responses comparable with computer responses; platforms differed significantly in attentiveness
Attentiveness, not ability scores
Two points follow from the table. First, the studies from 2022 to 2025 that compared MTurk and Prolific (Peer, Douglas, Uittenhove and Popovic and their colleagues) all put MTurk below Prolific on data quality, while the 2017 study found them comparable, and the two cognitive studies show the gap reaching ability scores, not only survey answers. Second, the size of the gap is a function of the screening rule. Popovic and colleagues report that even after stricter exclusions some test scores remained lower for MTurk participants, so exclusion narrows the gap without necessarily closing it.
3 Why Do Platform Rankings Keep Changing?
A platform ranking is a dated measurement of a changing population, so a result from one year cannot be carried to the next without a check. Peer and colleagues found in 2017 that Prolific Academic produced data quality comparable to MTurk's and more naive participants. By the 2022 study, MTurk quality was poor even when approval-rating filters were applied, while Prolific performed well without them. The authors propose that researchers assess data quality on an ongoing basis because platforms and panels change.
What changes underneath is partly who uses the site and why. In the first Peer and colleagues 2022 study, 68.7 percent of Prolific participants passed both attention checks, against 46.6 percent on MTurk, 22.5 percent on Qualtrics and 22.1 percent on Dynata. In that study, 41 to 42 percent of MTurk participants said they used the site as their main source of income, against 4 to 8 percent on Prolific and 12 percent on CloudResearch. The authors report that the lowest data quality came from MTurk participants who used the site as their main income source but spent few hours on it per week, and that usage patterns predicted quality better than reputation scores did. For a researcher, the practical reading is that a participant's relationship to the platform predicts quality better than the platform's label.
Disagreement about the ceiling of the problem is open. Cuskley and Sulik (2024, Perspectives on Psychological Science) argue that high-quality data is the responsibility of the researcher, not the platform, and that filtering services such as CloudResearch can bring MTurk data to a quality similar to, and sometimes higher than, Prolific's. They also write that Prolific is not immune to data-quality problems. Their recommendations are about design: fair pay, short surveys, piloting before recruitment and multistage filtering. Peer and colleagues' 2022 result, that MTurk stayed poor with filters, and Cuskley and Sulik's argument, that filters can rescue it, are not reconciled by either paper. A researcher who needs a defensible answer should run a short pilot on the platform, with the screening they plan to use, and read the pass rates before committing the sample.
The Douglas and colleagues (2023) comparison widens the picture to five sources. Across 2,729 respondents from MTurk, Prolific, CloudResearch, Qualtrics and SONA, the authors report that Prolific and CloudResearch participants were more likely to pass attention checks, provide meaningful answers, follow instructions, remember previously presented information, have a unique IP address and geolocation, and work slowly enough to read all the items. On a recall question about a video, 52.20 percent of MTurk participants answered correctly, against 83.47 percent on Prolific and 81.78 percent on CloudResearch. Those percentages describe a survey experiment, not a cognitive battery, but the checks they use, identity, reading time and memory for instructions, are the same ones an online cognitive study has to pass before the ability scores can be read.
4 Are Bots the Problem, or Are People?
The strongest recent evidence says most bad online data comes from inattentive or fraudulent humans, not from software, but the figures come from different samples and the dispute is not closed. The claim that online panels are overrun by bots has a vivid source. Webb and Tangney (2024, Perspectives on Psychological Science) describe running an MTurk study as a doctoral student and deducing that, at best, they had gathered valid data from 14 human beings, 2.6 percent of a sample of 529. The abstract frames this as the author's own experience and a call for caution, and notes that in 2015 an estimated 45 percent of articles in top behavioral and social science journals included at least one MTurk study.
Cuskley and Sulik answer in the same journal issue. They call the 2.6 percent figure "radically misleading" and say it is "off by more than an order of magnitude relative to even the most conservative estimates," attributing the low quality to the study's design choices rather than to MTurk as such. They cite research that about 60 percent of MTurk workers give acceptable data, and compare the 10 percent failure rate on instructional manipulation checks in the Webb and Tangney study with the 14 to 18 percent that they cite for motivated participants tested in person. We are reporting their argument, not settling it: the two papers use different samples, designs and definitions of a valid response, and neither can be applied to a cognitive battery without a pilot.
A third paper changes the question rather than picking a side. Jaffe and colleagues (2026, Perspectives on Psychological Science) examined where problematic data on online panels comes from, using four studies spanning five years and several platforms, including MTurk and Lucid. They report evidence that most of the data-quality problems affecting online research with online panels can be tied to fraudulent users from outside the United States, not to bots, and that these humans leave identifiable signs. The authors also describe the most effective ways of blocking such respondents.
For a researcher the label matters because the remedy differs. If the problem were software, a challenge that machines fail would be the control. If the problem is people who misrepresent where they are or who they are in order to collect the incentive, the control is verifying eligibility and location at the platform level and checking for duplicate or implausible identities in the data. Our reasoning is that both controls belong in a protocol, and that neither replaces looking at the actual response patterns of your own sample, which is where the next two sections turn. This page describes what to verify, not how any control is evaded.
5 What Do Attention Checks Catch, and What Do They Miss?
Attention checks catch participants who do not read, but in our reading they say little about effort on an ability task, and a participant who passes a common check can still give unusable data. The foundation paper is Meade and Craig (2012, Psychological Methods). In two studies they compared five approaches to identifying careless responses: special items designed to detect careless response, response consistency indices formed from typical survey items, multivariate outlier analysis, response time and self-reported diligence. They found two distinct patterns of careless response, random and nonrandom, and that different indices are needed to identify each. About 10 percent to 12 percent of the undergraduates who completed a lengthy survey for course credit were identified as careless responders. Their recommendations include using identified rather than anonymous responses, placing instructed response items before data collection, and computing consistency indices and multivariate outlier analysis.
Three details from the later literature qualify how far those checks go. First, the pass rate on a check depends on the pool: in the first Peer and colleagues 2022 study, 68.7 percent of Prolific participants passed both attention checks against 46.6 percent of MTurk participants. Second, checks vary in difficulty. Esch, Mylonopoulos and Theoharakis (2025) raise a concern that most MTurk users can pass frequently used attention checks but fail less utilized measures such as the infrequency scale, so a pool that looks clean on a familiar check may not be. Third, in the same study, attentiveness was affected by the respondent's incentive for completing the survey, their activity before engaging, environmental distractions and having recently completed a similar study. Those are conditions of the session, not traits of the person, and they apply to a cognitive task as much as to a survey.
Checks also differ in what they test. An instructed response item, such as an item telling the participant which option to choose, tests whether the text was read. A recall question tests whether instructions were retained. A consistency index tests whether responses contradict each other. None tests whether the participant tried to solve a hard item. For cognitive tasks the closest equivalents are benchmark effects the task is known to produce. In the Uittenhove study, the effect of irrelevant speech on rehearsal was missing from 42 percent of MTurk participants, against 16.9 percent on Prolific, and that was one criterion among several used to decide which datasets to keep. A benchmark of this kind needs a task with a documented, robust effect, so it is rarely available for a general ability battery.
The practical consequence is a layered rule rather than a single check: an instructed item before the task, a consistency or pattern index afterward, and a benchmark or catch trial if the task offers one. Whatever the combination, run it on a pilot first and look at how many participants each layer removes, because a rule that removes 60 percent of one group and 20 percent of another has changed the sample, not only cleaned it.
6 Does Excluding Careless Participants Bias the Sample?
Exclusion can protect an analysis and can also distort it, so the rule has to be justified by what it removes, not only by what it keeps. The protective case is strong for relational analyses. Zorowitz and colleagues (2023, Nature Human Behaviour) showed in two online samples (total N = 779) that careless responders can create spurious associations between task behavior and symptom scores, because many symptom surveys have asymmetric score distributions and careless responders therefore appear to have elevated symptoms. They report that false-positive rates for these spurious correlations increase with sample size, contrary to common assumptions, that excluding participants flagged for careless responding on surveys abolished the spurious correlations, and that exclusion based on task performance alone was less effective. The lesson for an ability study is that the more variables you correlate with a cognitive score, the more a small share of careless responders can mislead, and that a flag set independently of the outcome works better than one built from the outcome itself.
The distorting case is the cost side. When the exclusion rate differs by group, the groups are no longer comparable. Popovic and colleagues (2025, Scientific Reports) tested three face identity ability measures and report that MTurk participants scored approximately 10 percentage points lower than those recruited through Prolific and university students tested in the laboratory. Applying stricter exclusion criteria based on attention checks excluded about 62 percent of the MTurk group and about 22 percent of the Prolific group, and some test scores remained lower for MTurk participants even after exclusion. The same paper notes that the face tests had been developed using MTurk participants, so it provides updated normative scores. A norm built from a pool with high exclusion and lower scores is a different norm from one built on a pool with low exclusion (the page on how IQ scores are normed explains why the reference group defines the score), which is why the platform belongs in the methods section.
Three cautions follow, and all three are our reasoning rather than a finding in one paper. A high exclusion rate is a statement about attention and compliance in the recruited pool, not about the ability of the people removed, and nothing in the papers above supports reading it as low intelligence. Excluding on low scores is circular for an ability study, because it removes the lower tail and shrinks the variance you are trying to measure, which is why rules should be built from process indicators, such as response patterns, timing and instruction checks, and not from the score. And when the final sample is a minority of those recruited, as with the 34.7 percent of MTurk participants who passed every criterion in the working memory study, the survivors are a selected subset, so the sample describes attentive MTurk participants, not MTurk participants in general.
The table organizes the common rule families with their evidence and their cost. It is our summary of the sources on this page, and the cost column is reasoning, not a figure from any paper.
Rule family
What it targets
What the sources document
Cost or risk
Instructed response item
Not reading the text
Recommended by Meade and Craig (2012) before data collection; pass rates differ by pool (Peer and colleagues, 2022)
Passing it does not show effort on a hard item
Infrequency or rarely used check
Inattentive responding that passes familiar checks
Most MTurk users passed frequently used checks and failed the infrequency scale (Esch and colleagues, 2025)
Harder to explain to participants; needs a pilot
Consistency index and multivariate outliers
Random and nonrandom careless patterns
Different indices detect different patterns (Meade and Craig, 2012)
Needs enough items; unusual honest profiles can be flagged
Response time floor
Responding faster than the material can be read
Douglas and colleagues (2023) report who worked slowly enough to read all items; ITC and ATP guideline 4.26 on accurate timing
Meaningless on speeded subtests where fast is the signal
Benchmark effect or catch trial
Data that fail to show a robust known effect
42 percent of MTurk against 16.9 percent of Prolific lacked the rehearsal disruption effect (Uittenhove and colleagues, 2023)
Unique IP address and geolocation differed by platform (Douglas and colleagues, 2023); fraud from outside the United States (Jaffe and colleagues, 2026)
Shared networks can produce false duplicates
Exclusion on task score
Low performance
Less effective than survey based flags (Zorowitz and colleagues, 2023)
Circular for ability; shrinks variance
7 How Do Response Times and Speeding Rules Work in a Cognitive Study?
Response time is the cheapest process indicator a computer records, but a fast response is a defect on an untimed item and the signal itself on a speeded subtest, so the rule has to be set item by item. Meade and Craig list response time among their five approaches to detecting careless responding, and Douglas and colleagues count working slowly enough to read all the items as one of their data quality indicators, with Prolific and CloudResearch participants doing better than the others. For a reading-heavy item a plausible floor is the time it takes to read it, and a response below that floor is evidence that the item was not processed.
For ability items the floor has a second use, because the participant may be answering without trying. That is the reasoning behind a threshold rule, and the ITC and ATP guidelines give the measurement requirement that makes it usable. Guideline 4.26 says response times should be recorded as accurately as possible, avoiding the effects of technology requirements to respond and distractions in the environment, and measured with a high degree of precision; its comment says milliseconds are preferred over seconds and that individual differences in response times due to construct-irrelevant differences in computer technology should be avoided. Guideline 4.25 adds that if response times are used in scoring, this should be disclosed. Guideline 4.27 asks that construct-irrelevant factors that may affect response time, such as motor disabilities, executive function challenges and testing in a non-dominant language, should not negatively impact test takers' scores. A speeding rule built without that last point can remove a group of participants for reasons unrelated to effort.
Timed tasks need a different rule. On a speeded subtest such as symbol matching, speed is the quantity being measured and a fast response is the intended behavior, so the screening question becomes whether the response times are plausible given the input method and the device, not whether they are too short. The page on processing speed explains why the device is part of the measurement for speeded tasks, and the page on why IQ tests are timed explains what the clock measures. The next section gives the evidence on how devices shift reaction times.
Timing also helps against assistance. Steger, Schroeders and Wilhelm (2021, Assessment) had 315 participants complete a knowledge test first in an unproctored online assessment and then in a proctored laboratory session. They compared three kinds of indicators for predicting who cheated: self-report scales (Honesty-Humility and Overclaiming), test data (performance with extremely difficult items) and para data (reaction times and switching between browser tabs). Test data and para data performed best, while the traditional self-report indicators were not predictive. That was a knowledge test, not an intelligence battery, so the result transfers to ability testing only as a direction for pilot work. The direction is useful: what a participant did during the task, recorded unobtrusively, predicted more than what the participant said about themselves.
8 How Do Device and Environment Change Scores on Timed Tasks?
The device is part of the measurement on any task that records reaction time, and a comparison of 59,587 participants found mobile devices, particularly Android smartphones, slower on reaction time tests than laptops and desktops. Passell and colleagues (2021, Behavior Research Methods) analyzed cognitive test data from 59,587 participants in their first study and 3,818 in their second. They report that users of mobile devices, particularly Android smartphones, showed significantly slower performance on tests of reaction time than users of laptop and desktop computers. The title of the paper states the general point: cognitive test scores vary with the choice of personal digital device.
The finding bears on a screening decision, not only on interpretation. If a participant on a phone produces slower reaction times than the same person would on a desktop, then a time-based exclusion or a processing speed score will mix ability with hardware. The ITC and ATP guidelines recognize the option of constraining the device. The comment to guideline 3.21(a) says that "the allowed classes of devices could be limited to offer a comparable testing condition to all end users," by listing allowed or disallowed device classes, form factors, input types such as keyboard, mouse or touchscreen, screen orientation, operating systems and browser versions. Uittenhove and colleagues took that route: they ensured participants used a laptop or desktop rather than a mobile device. The cost of the restriction is coverage, since people who mainly use phones are excluded by design.
Mobile access is not necessarily a threat to attention. In the Esch, Mylonopoulos and Theoharakis study, most respondents on the Pollfish and Qualtrics panels were mobile-based while most on MTurk and Prolific were computer-based, and using an attentiveness composite the authors found mobile-based responses comparable with computer-based responses. The same paper found attentiveness affected by environmental distractions and recent activity. Taken together, the two papers separate two questions: whether a phone user is attentive, which the Esch result answers favorably, and whether a phone produces comparable timing, which the Passell result answers unfavorably for reaction time measures. A study of reasoning accuracy with generous time limits may tolerate mixed devices where a study of speeded tasks cannot.
The environment is the other half. The Standards for Educational and Psychological Testing say in Standard 6.4 that the testing environment should furnish reasonable comfort with minimal distractions to avoid construct-irrelevant variance, and the comment names noise, disruption, poor lighting and malfunctioning computers among the conditions to avoid. An unsupervised participant controls none of these for the researcher. Standard 6.3 asks that changes or disruptions to standardized administration be documented and reported to the test user, with the comment noting that a researcher may want to use only records based on standardized procedures. For an online study, the equivalent is to record what the platform allows about device, browser and input method, and to ask participants about their environment, so that the analysis can compare conditions rather than assume them. For the domain behind the working memory task in the Uittenhove study, see the page on working memory testing for adults.
9 How Much Do Unproctored Ability Tests Inflate Scores?
Across 49 studies, unproctored ability assessments produced higher scores than proctored ones by a pooled standardized difference of 0.20, which the authors read as evidence of cheating, with a smaller effect for tasks that are hard to look up. Steger, Schroeders and Gnambs (2020, European Journal of Psychological Assessment) ran a three-level random-effects meta-analysis of 109 effect sizes from 49 studies, with a total N of 100,434. The pooled effect was Δ = 0.20, 95 percent confidence interval 0.10 to 0.31, higher in the unproctored condition. The moderators they examined were perceived consequences of the assessment, countermeasures against cheating, susceptibility of the measure itself and the test medium. Of those, the abstract reports one result: significantly smaller effects for measures that are difficult to research on the Internet. They conclude that unproctored ability assessments are biased by cheating, and suggest unproctored assessment may be most suitable for tasks that are difficult to search on the Internet.
A pooled difference of 0.20 is small against a standard deviation of 15 IQ points, and by our arithmetic it corresponds to about 3 points on that scale if the metric of the studies were translated directly, which the meta-analysis does not do. The distribution matters more than the mean: a pooled difference can come from a few participants gaining a lot or from many gaining a little. For group comparisons the mean shift is a bias only if the conditions differ between groups, so the risk is greatest when one group is tested unproctored and another is tested in a laboratory.
The ITC and ATP guidelines place these risks in a broader frame. Chapter 8 lists, among the categories of score validity threats due to cheating, using pre-knowledge about the test, receiving expert help while taking the test, using unauthorized test aids or assistance, using a proxy test taker, tampering with testing software or stored test results, copying answers from another test taker and manipulating testing rules. Guideline 8.7 asks testing organizations to detect and report cheating, naming data forensics, monitoring Internet sources for disclosed content and monitoring the test taker as possible measures. The comment adds that when initial results are taken under non-secure conditions, such as no proctor at all, a verification test under secure conditions provides an opportunity to verify whether the initial exams were taken without cheating. In our reading that is a design option for a research study too: a short proctored or live follow-up on a subsample could calibrate the unproctored data.
Two qualifications keep this evidence in proportion. The meta-analysis includes assessments with consequences for the test taker, and the authors list perceived consequences as a moderator, so, in our reading, a paid research participant with no stake in the score is in a different position from a job applicant. And the Steger and colleagues 2021 study found that the indicators that best predicted cheating were test data and para data, which are recorded without asking the participant anything, as the section on response times describes.
10 What Do Repeat Takers and Circulating Content Do to a Sample?
Participants who have seen similar tasks, or the items themselves, are a second source of non-independence, and the literature measures it through naivety and recent task completion more than through direct counts of repeaters. Peer and colleagues (2017) found that both alternative platforms they tested, CrowdFlower and Prolific Academic, attracted more naive participants than MTurk, which is a statement about prior exposure to research tasks. In the Esch and colleagues study, having recently completed a similar study was one of the factors that affected attentiveness. Douglas and colleagues report that Prolific and CloudResearch participants were more likely to have a unique IP address and geolocation, which is a platform-level indicator that one person was not entered more than once.
The ITC and ATP guidelines describe the content side of the same risk in their list of threats due to test content theft: stealing test files, stealing questions using digital photography or electronic capture, memorizing content for later recording or sharing, transcribing questions into a recording device and obtaining test material from a trusted insider. The Standards make the fairness point from the other direction: unauthorized distribution of items to some examinees but not others could provide an advantage to some test takers, and in that case results should be interpreted with caution. This page lists the risks because a researcher has to know what to ask a test publisher (the page on free versus validated IQ tests gives a buyer's version of the same questions); it does not describe how any of them is carried out.
What follows for design is mostly about the instrument and the recruitment. A task whose items are widely published, such as a short problem set that appears in textbooks and social media, will have more prior exposure in any large online pool than an unpublished one, and an item bank that is public cannot promise that an item is unseen. The page on the Cognitive Reflection Test deals with a task of this kind, and the page on the ICAR item pool covers how a public bank handles exposure. On the recruitment side, a protocol can ask participants about prior experience with the task. A participant who says they have seen the task before is a data point for a sensitivity analysis, not automatically an exclusion.
The size of practice effects on a given cognitive task is a question for that task's own retest data, and this page gives no figure for it because the sources above did not measure it. The relevant rule from the guidelines is disclosure: say what the participant pool had been exposed to, and say what you did about it.
11 How Should You Read a High Exclusion Rate, and What Should You Write Down in Advance?
A high exclusion rate is a finding about the recruited pool and the screening rule, and it should be reported with the rule that produced it, preferably one written down before the data were seen. The studies above show that rates vary widely with the pool and the rule: about 62 percent against about 22 percent in the face recognition study with stricter attention checks, and 65.3 percent against 28.7 percent failing at least one criterion in the working memory study, by our subtraction from the 34.7 and 71.3 percent that passed everything. Neither pair of figures is a benchmark for your study. Each is an example of how far the same screening can differ across pools.
A defensible protocol is a short list that turns those findings into steps. It is our synthesis, and each step points back to a source above.
Pilot first. Run the intended recruitment filters and screening on a small sample and read the pass rate by rule, as Peer and colleagues recommend ongoing assessment of platform quality.
Write the rules down before the main data arrive. Name each rule, its threshold and the order in which it applies, so that the decisions cannot follow the results.
Build rules from process, not score. Use instruction checks, pattern indices, timing and duplicate checks, and avoid excluding on the ability score itself, for the reasons Zorowitz and colleagues report.
Apply the same rules to every group. If groups differ in exclusion rate, report that, since it is itself a result.
Record the conditions. Log platform, date, filters, compensation, device class, input method and the mode of supervision, in line with the ITC and ATP guidance on interpreting outcomes in light of conditions and with Standard 6.3 on documenting disruptions.
Count what each rule removed. Report the number excluded by each rule, not only the total.
Run the analysis both ways. Report results with and without the exclusions, as a sensitivity check, so that readers can see whether the conclusion depends on the rule.
ACIS is an online, unsupervised assessment, so everything this page says about unproctored data applies to it, and a buyer should read its administration and documentation pages before relying on any score. The facts below were read on the ACIS home page on October 6, 2026, and prices can change. ACIS has 20 subtests in six domains of the CHC model and reports a Full Scale IQ and six primary indices: verbal comprehension, fluid reasoning, quantitative reasoning, visual spatial, working memory and processing speed. There are three forms, each a one time payment with no subscription: Quick at 15 dollars with 6 subtests in 3 domains and about 45 minutes, Optimized at 30 dollars with 13 subtests in 5 domains and about 110 minutes, and Full Scale at 50 dollars with all 20 subtests in six domains and about 175 minutes. Breaks are allowed, there is a free trial with no card, a 5 day quality guarantee and 30 days to complete a form. Indices and the Full Scale IQ are reported on the standard scale with a mean of 100 and a standard deviation of 15, subtests as scaled scores from 1 to 19, and the report gives percentiles and a 95 percent confidence interval. Adult norms cover ages 16 to 90.
The limits are plain. ACIS is online and unsupervised, it is not a clinical or diagnostic instrument, it is not for hiring, school accommodations or admission to high IQ societies, and it is available in English only. This page reports no ACIS exclusion rate and says nothing about how ACIS handles the quality of its own data, because that is not what it documents here. A researcher who wants to know how administration works for a study should read the page on ACIS for research, and anyone who needs the documentation behind the scores should read the technical manual rather than rely on a summary from a third party, including this page.
What a buyer should do is the same for every provider. The Standards for Educational and Psychological Testing (AERA, APA and NCME, 2014), co-published by the American Psychological Association, say that the interpreter of the results should be informed if a test was unproctored, that unproctored administrations where standardization cannot be ensured may advantage some test takers, and that results in such situations should be interpreted with caution. The ITC and ATP guidelines ask for guidance on test-taking conditions, on the level of authentication and on monitoring in proportion to the stakes. Applied to a purchase (the page on IQ test apps does this for consumer products), that means asking any provider, ACIS included, four questions: what conditions the test assumes, how identity and device are handled, what the score report states about the administration, and where the evidence is documented. The pages on online versus professional IQ tests and types of IQ tests set out the supervised and unsupervised split for individual test takers, and the page on what to verify before you sign with a platform does the same for organizations.
Every figure on this page comes from the documents below, opened or verified on October 6, 2026, and every figure we derived ourselves is labeled in the text as our arithmetic or reasoning.
American Educational Research Association, American Psychological Association, and National Council on Measurement in Education (2014). Standards for Educational and Psychological Testing. Washington, DC: AERA. Open access edition: https://www.testingstandards.net/open-access-files.html
Meade, A. W., and Craig, S. B. (2012). Identifying careless responses in survey data. Psychological Methods, 17(3), 437 to 455. https://doi.org/10.1037/a0028085
Peer, E., Brandimarte, L., Samat, S., and Acquisti, A. (2017). Beyond the Turk: Alternative platforms for crowdsourcing behavioral research. Journal of Experimental Social Psychology, 70, 153 to 163. https://doi.org/10.1016/j.jesp.2017.01.006
Peer, E., Rothschild, D., Gordon, A., Evernden, Z., and Damer, E. (2022). Data quality of platforms and panels for online behavioral research. Behavior Research Methods, 54(4), 1643 to 1662 (published online 2021). https://doi.org/10.3758/s13428-021-01694-3
Douglas, B. D., Ewell, P. J., and Brauer, M. (2023). Data quality in online human-subjects research: Comparisons between MTurk, Prolific, CloudResearch, Qualtrics, and SONA. PLOS ONE, 18(3), e0279720. https://doi.org/10.1371/journal.pone.0279720
Uittenhove, K., Jeanneret, S., and Vergauwe, E. (2023). From lab-testing to web-testing in cognitive research: Who you test is more important than how you test. Journal of Cognition, 6(1). https://doi.org/10.5334/joc.259
Popovic, B., Dunn, J. D., Towler, A., and White, D. (2025). Normative face recognition ability test scores vary across online participant pools. Scientific Reports, 15(1). https://doi.org/10.1038/s41598-025-92907-8
Esch, D., Mylonopoulos, N., and Theoharakis, V. (2025). Evaluating mobile-based data collection for crowdsourcing behavioral research. Behavior Research Methods, 57(4). https://doi.org/10.3758/s13428-025-02618-1
Webb, M. A., and Tangney, J. P. (2024). Too good to be true: Bots and bad data from Mechanical Turk. Perspectives on Psychological Science, 19(6), 887 to 890. https://doi.org/10.1177/17456916221120027
Cuskley, C., and Sulik, J. (2024). The burden for high-quality online data collection lies with researchers, not recruitment platforms. Perspectives on Psychological Science, 19(6), 891 to 899. https://doi.org/10.1177/17456916241242734
Jaffe, S. N., Moss, A. J., Hartman, R., Rosenzweig, C., Gautam, R., Robinson, J., and Litman, L. (2026). The bots ruining social science are not bots at all. Perspectives on Psychological Science, 21(2), 127 to 137. https://doi.org/10.1177/17456916251404872
Zorowitz, S., Solis, J., Niv, Y., and Bennett, D. (2023). Inattentive responding can induce spurious associations between task behaviour and symptom measures. Nature Human Behaviour, 7(10), 1667 to 1681. https://doi.org/10.1038/s41562-023-01640-7
Passell, E., Strong, R. W., Rutter, L. A., Kim, H., Scheuer, L., Martini, P., Grinspoon, L., and Germine, L. (2021). Cognitive test scores vary with choice of personal digital device. Behavior Research Methods, 53(6), 2544 to 2557. https://doi.org/10.3758/s13428-021-01597-3
Steger, D., Schroeders, U., and Gnambs, T. (2020). A meta-analysis of test scores in proctored and unproctored ability assessments. European Journal of Psychological Assessment, 36(1), 174 to 184. https://doi.org/10.1027/1015-5759/a000494
Steger, D., Schroeders, U., and Wilhelm, O. (2021). Caught in the act: Predicting cheating in unproctored knowledge assessment. Assessment, 28(3), 1004 to 1017. https://doi.org/10.1177/1073191120914970
14 Frequently Asked Questions
What does data quality mean in online cognitive testing?
It means whether the scores reflect each participant's ability under comparable conditions. Quality can fail through the wrong person, low effort, an unsuitable device, outside help or prior exposure to the items. Each failure leaves a different trace, so screening needs several indicators rather than one number.
Is Prolific better than MTurk for cognitive research?
In the studies reviewed here, Prolific data were closer to laboratory and web-tested student data than MTurk data were, including on a working memory task. Rankings have changed over time, and filtering services have been argued to narrow the gap, so a pilot on your own task is safer than a platform label.
How common is careless responding in online studies?
It depends on the sample and the definition. In one early study, about 10 percent to 12 percent of undergraduates completing a lengthy survey for course credit were identified as careless. Pass rates on attention checks differed by more than 20 percentage points across platforms in later work.
Do bots ruin online cognitive research?
The evidence is divided. One researcher described a sample that was at best 2.6 percent valid, while other authors called that figure misleading and argued that problems come mainly from careless or fraudulent humans. Either way, a study should verify eligibility and check its own response patterns.
Are unproctored online cognitive tests valid?
They can support some uses and not others, depending on the stakes and the task. A meta-analysis found higher scores without a proctor, with smaller effects on tasks that are hard to look up. The test standards ask that interpreters be told when a test was unproctored.
What is an attention check and does it work?
It is an item that tests whether a participant read or followed instructions, such as an instructed response item. It catches non-reading but not necessarily low effort on a hard task. Some pools pass familiar checks and fail rarer ones, so layered indicators work better than one check.
How do I screen online cognitive data?
Combine process indicators that do not depend on the ability score: instructed items, response pattern indices, plausible timing, duplicate and location checks and, where a task offers one, a known benchmark effect. Pilot the rules, write them down before the main data arrive and report what each removed.
What did Uittenhove and colleagues find about MTurk and Prolific?
Using a working memory task with 751 participants, they found that 71.3 percent of Prolific and 34.7 percent of MTurk participants passed all quality criteria, against 72.6 percent of web-tested students and 87.6 percent of laboratory students. They concluded that who you test matters more than how.
What did Peer and colleagues find in 2022?
In two studies totaling about 4,000 participants, Prolific provided high quality on attention, comprehension, honesty and reliability measures, CloudResearch did well when filters were applied, and MTurk performed poorly even with filters. Usage patterns, such as relying on the site for income, predicted quality better than reputation scores.
How much higher are unproctored ability scores?
A meta-analysis of 49 studies found a pooled standardized difference of 0.20, with a 95 percent confidence interval of 0.10 to 0.31, favoring unproctored conditions. The authors reported smaller effects for measures that are difficult to look up on the Internet and read the result as evidence of cheating.
Do phones and tablets change cognitive test scores?
For reaction time, yes in one large comparison: users of mobile devices, particularly Android smartphones, were slower than laptop and desktop users, in samples of 59,587 and 3,818 participants. A separate study found mobile responses comparable with computer responses on an attentiveness composite, so attention and timing are separate questions.
What did the face recognition study find about exclusion?
Applying stricter attention-check criteria excluded about 62 percent of MTurk participants and about 22 percent of Prolific participants. Before exclusion, MTurk scored approximately 10 percentage points lower on all three tests, and some scores stayed lower afterward. The authors therefore published updated normative scores.
What do the ITC and ATP guidelines require for unproctored tests?
The 2022 guidelines ask that testing conditions be appropriate and outcomes interpreted in light of them, that guidance cover the required level of identity authentication, that procedures be monitored in proportion to stakes, and that response times be recorded accurately. They also ask organizations to detect and report cheating.
What do the Standards say about unproctored administration?
They say the interpreter of results should be informed if a test was unproctored or given under nonstandardized procedures, and that unproctored administrations where standardization cannot be ensured could advantage some test takers, so results should be interpreted with caution. They also treat the mode of administration as a validity consideration.
Should I exclude participants who fail an attention check?
Usually yes if the rule was set in advance, applied equally to all groups and reported, and then check whether conclusions change when those participants are kept. Careless responders can create false correlations, but a rule that removes very different shares of different groups changes what the sample represents.
Does a high exclusion rate mean participants had low ability?
No. An exclusion rate describes attention, compliance and fit with the screening rule in the pool you recruited, not the ability of the people removed. Excluding on a low ability score is a different and circular rule, and the sources reviewed here do not support reading exclusions as low intelligence.
How many participants should I recruit to allow for exclusions?
Divide the sample you need by the pass rate you expect from a pilot, then add margin. As our arithmetic, a pass rate of 34.7 percent implies recruiting about 2.9 people for every usable dataset, while 71.3 percent implies about 1.4. Pass rates vary by task and pool.
Should I preregister exclusion rules?
Yes, wherever you can. Naming each rule, its threshold and its order before seeing the outcome data stops the decisions from following the results. If a rule must change after a pilot or a problem, record what changed and why, and report the main findings both with and without exclusions.
What should I report about data quality in a paper?
Report the platform and dates, recruitment filters, compensation, device and input restrictions, whether the session was supervised, each exclusion rule with its threshold, the number removed by each rule, and the results with and without exclusions. Name the instrument and administration mode as well.
Can I restrict a study to desktop computers?
Yes, and one cognitive study did so, requiring laptops or desktops. The ITC and ATP guidelines note that allowed device classes can be limited to give comparable conditions. The cost is coverage, since people who mainly use phones are excluded, which should be stated as a limit.
Can I use an unsupervised online IQ test in a research study?
Yes, where the design can tolerate unsupervised conditions and the limits are reported. ACIS, for example, is online and unsupervised and is not a clinical or diagnostic instrument. Pilot your screening, state the mode of administration and avoid using the scores for decisions that need a proctored result.
Take the assessment
You get a profile, not a number
ACIS measures six CHC domains across 20 subtests and reports each one with its own normed score and confidence interval, so you can see where you are strong and where you are not.