Cognitive Assessment Platforms, Sorted by What They Actually Answer
The phrase cognitive assessment platform covers four unrelated product categories: clinical screeners, research task builders, consumer brain training and normed ability batteries. Every ranking listicle mixes them. Buying from the wrong category produces data that cannot answer your question, no matter how good the product is. Separate the categories first.
Ranking listicles put clinical screeners, experiment builders and normed ability batteries in one list. They are four different purchases with four different outputs.
0 Quick Answer: One Phrase, Four Different Products
A cognitive assessment platform is software for administering cognitive tests to people other than yourself, and the phrase currently names four product categories that do not compete with one another. The first is the clinical screener, built around a care protocol, a documented observation and a health record. The second is the research platform, where you either assemble your own tasks or administer a fixed research battery and take away trial level data. The third is consumer brain training, which optimizes for repeated engagement. The fourth is the normed ability battery, which places a person on a population scale and can return a composite in the IQ metric.
The complication is that these four sit in the same search results and the same ranking articles. A list titled "best cognitive assessment tools" will happily rank Creyos, Gorilla and a brain training app in one table, scored on one rubric, as if a buyer could substitute one for another. They cannot. A clinical screener does not produce a research dataset. An experiment builder does not produce a normed score. A normed battery does not produce a billable clinical observation. Each is excellent inside its category and useless outside it.
The practical consequence is money spent on a product that cannot answer the question that justified the spend. A study team that buys a clinical screener discovers it cannot export the trial level timing it needed. A screening program that buys an experiment builder discovers it has raw accuracy counts and no reference frame to interpret them against. Both purchases were competent products bought for the wrong job.
4
Unrelated product categories that share the phrase cognitive assessment platform.
1 of 4
Categories that returns a composite in the IQ metric at all.
44,600
Adults in the Hampshire, Highfield, Parkin and Owen 2012 Neuron analysis that split one widely used task battery into three components rather than one general score.
Read this firstDecide your category before you compare vendors. A rubric that ranks across categories is measuring products against a job most of them were never built to do, which makes the ranking meaningless even when every individual fact in it is true.
1 The Four Categories, Defined by What They Output
Categories are easiest to separate by their output, not by their marketing. Marketing language converges: every vendor in every category says validated, objective, digital and scalable. Output does not converge. Ask what artifact lands at the end of an administration and the four categories fall apart cleanly.
Clinical and care pathway screeners output a documented observation inside a patient record. The score matters less than the fact that it is timestamped, attributable, comparable with the same patient's earlier result and structured so a clinician can act on it and a payer can process it. Creyos is the clearest example. Its materials describe a cognitive assessment platform for healthcare, and its results are described as coded for reimbursement and built for the electronic health record.
Research platforms output trial level data. This category splits into two sub types that buyers confuse more often than they confuse the four main categories. A task builder such as Gorilla, Testable or PsyToolkit gives you an authoring environment: you assemble the paradigm, you own the design, and nothing about the resulting numbers is normed. A fixed research battery such as the NIH Mobile Toolbox or CANTAB gives you published instruments you administer as designed, with published psychometrics for each measure, and usually no single composite.
Consumer brain training outputs engagement and a progress score inside the product. It exists to be used repeatedly by the same person. The number it shows is a within product training metric, and it is not built to place someone in a population.
Normed ability batteries output a standardized score on a population scale, typically scaled subtest scores and composites that include something in the IQ metric. Pearson's Q-interactive is the supervised, publisher grade version of this. ACIS Professional is the unsupervised, self serve version. What both share, and what nothing in the other three categories offers, is a reference frame that makes one person's number interpretable on its own.
The output determines everything downstream: who is permitted to administer, whether an individual score can be interpreted, whether the data will support a group comparison, and whether anything can be exported. Read what IQ scores mean for how an interpretable individual score differs from a raw performance count, and the six cognitive domains for what a multi domain profile contains.
2 Route From Your Question to the Right Category
Write down the sentence you want to be able to say at the end of the project, then read this table backwards from that sentence. Most bad platform purchases are made by teams who never wrote the sentence, compared feature lists instead, and picked the product with the longest one.
Your question
Category you need
Platforms in that category
What it will not tell you
Is this patient showing cognitive change I should document, act on and bill for?
Clinical screener
Creyos, CANTAB clinical, Pearson Q-global
Where the person sits on a general population ability scale
Did my experimental manipulation move accuracy or reaction time?
Research task builder
Gorilla, Testable, PsyToolkit
Anything about an individual participant's standing, because there are no norms
How does my cohort compare with a published reference sample on established measures?
Fixed normed research battery
NIH Mobile Toolbox, CANTAB, Cognitron
A single overall score, since most report at the measure level by design
Do I need an individually administered result that a school, court or clinic will accept?
Supervised publisher battery, qualified examiner
Pearson Q-interactive running WAIS-5 or WISC-V
Nothing an unsupervised platform can substitute for, at any price
Do I want a normed multi domain profile with an IQ metric composite, administered remotely and at low cost?
Normed ability battery, online
ACIS Professional
A diagnosis, a proctored result, or a decision about hiring or accommodations
Do I want participants to practice tasks repeatedly and enjoy doing it?
Consumer brain training
Not an assessment purchase
Their ability level, since a training score is a within product metric
Two rows deserve emphasis. The fourth row is a hard boundary, not a preference: if the result must survive challenge by a third party, the administration model matters more than the instrument, and that is a qualification question rather than a software question. It is covered in full at who can administer an IQ test.
The third row is the one that catches research teams. Fixed normed batteries publish per measure psychometrics, so you can defend a claim about processing speed or working memory. They usually do not publish a composite, which means you cannot make a claim about general ability without doing your own factor work. If your hypothesis is stated at the level of general ability, you need a battery that models it, and the case for that is set out in the g factor explained.
3 Clinical Screeners: Creyos and the Care Protocol Model
Creyos describes itself as a cognitive assessment platform for healthcare, and every design decision follows from that word. Its published materials list twelve online cognitive tasks grouped into four areas: short term memory (Number Ladder, Spatial Span, Paired Associates, Token Search), reasoning (Odd One Out, Spatial Planning, Rotations, Polygons), concentration (Double Trouble, Feature Match) and verbal ability (Digit Span, Grammatical Reasoning). Alongside them sit standard behavioral health questionnaires including the PHQ-9, GAD-7, SWAN and AUDIT, plus condition focused protocols for ADHD and dementia screening.
The commercial logic is the health record, not the score. Creyos states that results are generated the moment an assessment ends, structured for documentation, coded for reimbursement and built for the electronic health record, with support for common CPT codes covering neuropsychological and cognitive services. That integration is the product. A research platform with better timing precision would still be the wrong purchase for a clinic, because it produces nothing a billing system can consume.
The task set has a documented research provenance worth understanding, because it explains why no IQ appears anywhere in the product. The same twelve task battery is the one analyzed by Hampshire, Highfield, Parkin and Owen in Fractionating Human Intelligence, published in Neuron in 2012, volume 76, pages 1225 to 1237. From roughly 110,000 people who logged in and 60,000 who finished all twelve tasks, exactly 44,600 complete datasets survived their exclusions. Performance decomposed into three components, short term memory, reasoning and verbal, rather than into one general factor. The authors proposed that intelligence is an emergent property of anatomically distinct systems.
That is a deliberate theoretical position, and a platform built on it will not hand you a single number. In the same paper a laboratory pilot of 35 participants produced a bivariate correlation of .65 between the mean standardized battery score and the Cattell Culture Fair test, which the authors reported as evidence of a relationship, not of equivalence. Creyos separately states a normative database of more than 85,000 participants drawn from a larger pool of 14 million test scores, and reports that the full twelve task assessment takes 35 to 45 minutes. What it does not publish is a technical manual, a reliability coefficient or a price, which leaves a $250 monthly starting figure on two software directories as the only public number, traced in the detailed review of Creyos against a vendor that lists its per administration cost.
What this category cannot doA clinical screener answers whether this person changed, or whether this person differs from a clinical benchmark. It is not built to answer where this person stands in the adult population on a general ability scale, and treating a domain percentile from a screener as an IQ estimate is a category error.
4 CANTAB: The Regulated Trial and Neuropsychiatric Battery
CANTAB, from Cambridge Cognition, occupies the space where cognitive measurement has to satisfy a regulator, a sponsor or an ethics committee. The company describes CANTAB as scientifically validated, highly sensitive, precise and objective measures of cognitive function correlated to neural networks, delivered on iPad or through web based testing, with complementary voice and questionnaire tools.
The measured domains are attention and psychomotor speed, executive function, memory, and emotion and social cognition. The condition coverage named on the vendor's own pages runs through ADHD, Alzheimer's disease, depression, Parkinson's disease and schizophrenia. The stated audiences are drug development companies running clinical trials, academic researchers and healthcare clinicians. Cambridge Cognition states on its cognitive tests page that CANTAB has been published in over 3,000 peer reviewed papers, a vendor claim that a buyer should treat as a claim rather than as an independently audited count, though the existence of a very large CANTAB literature is not in dispute.
What CANTAB is designed to detect is a difference: between a treated arm and a placebo arm, between a patient group and a matched control group, or between one visit and the next. Sensitivity to change is the engineering goal, and it is why the battery is heavy on tasks with adaptive difficulty and fine grained latency measurement rather than on knowledge based items.
No IQ appears in the CANTAB output, and it would be out of place if it did. A composite in the IQ metric compresses a broad profile into one highly reliable number precisely so that it can be compared across people. A trial endpoint needs the opposite property: a narrow, sensitive measure of one process that a compound might plausibly affect. These are competing design goals, and no battery optimizes both at once.
The buying implication is that CANTAB is rarely a self serve purchase. Configuration, licensing and support are structured for programs rather than for individual investigators buying a seat with a card. If your project is a single study with a modest budget, the fixed normed batteries in the next section and the task builders in the section after it will usually fit better, and the arithmetic behind that judgment lives at the per administration and per seat comparison. For how a research battery differs from a clinically administered intelligence scale, see professional IQ test versus online IQ test.
5 Research Task Builders: Gorilla, Testable and PsyToolkit
A task builder is an authoring environment, not a test. You design the paradigm, define the trials, set the timing, deploy the link and receive a spreadsheet of trial level responses. Nothing in the output is normed, and nothing in it interprets an individual. That is not a shortcoming, it is the category.
Gorilla is the most cited of the three. Its methods paper, Gorilla in our midst: An online behavioral experiment builder by Anwyl-Irvine, Massonnie, Flitton, Kirkham and Evershed, appeared in Behavior Research Methods in 2020, volume 52, pages 388 to 407. The validation is the honest part. In a first experiment with 268 children tested in supervised school settings in France and the United Kingdom, the flanker conflict effect appeared at 62.33 milliseconds. In a second experiment with 99 adults testing unsupervised at home, the same effect appeared at 29.1 milliseconds, statistically significant but smaller than the roughly 109 to 120 milliseconds typically reported from laboratory administrations. Online administration works, and it attenuates effects. Any team planning an unsupervised study should size their sample against the smaller number, not the laboratory one.
PsyToolkit takes the opposite commercial approach. Gijsbert Stoet described it in Teaching of Psychology in 2017 as a free web based service for setting up, running and analyzing online questionnaires and reaction time experiments. At the time of that paper the platform carried a library of more than 20 cognitive paradigms and more than 80 peer reviewed psychological scales, with documentation aimed at letting students work independently. For teaching and for small studies with no budget, it is difficult to beat.
Testable combines a visual experiment builder with a recruited participant pool, described by the company as Testable Minds, with identity verification before a study runs. That pairing matters for anyone whose bottleneck is finding participants rather than building the task.
62.33 ms
Flanker conflict effect in supervised school testing, Anwyl-Irvine et al. 2020, 268 children.
29.1 ms
The same effect measured unsupervised at home in the same paper, 99 adults.
No norms
What every task builder outputs at the individual level, by design.
The mistake this category invitesBuilding a Stroop task, a digit span and a matrix task in a builder does not produce a cognitive battery. It produces three raw scores with no reference frame, no evidence that they cohere, and no basis for saying a participant is above or below average. The reference frame is the expensive part, and it is described at how IQ scores are normed.
6 The NIH Mobile Toolbox: Published Psychometrics as the Benchmark
The Mobile Toolbox is the clearest public example of what a validated remote battery looks like when the evidence is actually published. The platform describes itself as bringing cognitive and other assessments to mobile devices for self administered remote testing across the adult lifespan, ages 18 to 90 and older, in English and Spanish on iOS and Android. Funding is stated as National Institute on Aging grant 2U2CAG060426. Its own summary claims more than 30 peer reviewed publications, more than six years of research and more than 3,000 study participants.
Two of those publications are worth reading before any purchase in any category, because they show the shape of real validation. In the Mobile Toolbox sequences task paper in Frontiers in Psychology, Slotkin, Kaat, Young, Dworak, Novack, Shono and colleagues reported in 2025, in volume 15, on a smartphone working memory measure across three samples: 92 in person, 1,007 remote and 147 in a test retest subsample. Split half reliability was high, with a median of .90. Convergent correlation with the NIH Toolbox List Sorting test was .64 in the first study and fell to .46 in the second. Two week test retest reliability was moderate, with an intraclass correlation of .55. Discriminant correlations with a picture vocabulary measure were appropriately lower.
The second is the validation of International Cognitive Ability Resource measures implemented in the Mobile Toolbox, published by Young, Kim, McKee, Doyle, Novack, Revelle, Gershon and Dworak in the Journal of Intelligence in 2025. In a sample of 100 United States adults aged 18 to 82, with 56 retested 11 to 17 days later, Puzzle Completion showed an alpha of .90 and test retest reliability of .82, correlating .40 with Raven's Progressive Matrices. Block Rotation showed an alpha of .93 and test retest reliability of .63, correlating .46 with a mental rotation test.
Read those numbers carefully, because they set a realistic bar. Internal consistency is high. Convergent correlations with the established reference instruments are moderate, in the .40 to .64 range, and they shrank when administration moved from supervised to remote. Test retest ranged from .55 to .82 depending on the measure. This is what honest remote measurement looks like, and any vendor claiming uniformly higher figures without publishing the studies deserves the same scrutiny. Compare against the framework at reliability and validity.
7 Cognitron: Research Provenance and a Commercial Present
Cognitron is the platform behind some of the largest remote cognitive datasets ever collected, and it is also now a commercial product, which are two different things a buyer must separate. The research system was developed by Adam Hampshire and Peter Hellyer at Imperial College London and used for the Great British Intelligence Test, run in association with a BBC2 Horizon collaboration. The public Great British Intelligence Test site describes a series of short tests of roughly one to four minutes each covering memory, reasoning, concentration and planning, with the full sequence and questionnaire taking about 40 to 50 minutes.
The scale it achieved is genuinely unusual. In Cognitive deficits in people who have recovered from COVID-19, Hampshire, Trender, Chamberlain, Jolly, Grant, Patrick, Mazibuko, Williams, Barnby, Hellyer and Mehta published in eClinicalMedicine in 2021, volume 39, article 101044, the analysis covered 81,337 participants who had completed both the cognitive tests and a questionnaire about suspected and confirmed infection. The largest effect appeared in hospitalized patients who had been ventilated, at 0.47 standard deviations of global cognitive deficit. The authors described that as greater than the average 10 year decline in global performance between ages 20 and 70, and equated it to about a 7 point difference on classic IQ tests.
That last phrase repays attention. Even a group with a very large dataset and a strong reason to communicate an effect size converted to IQ points for the reader rather than reporting an IQ. The platform measures performance and models group differences. It does not place an individual on a normed intelligence scale, and the 7 point figure is a translation of a group mean difference, not a score anyone received. Group statistics never license an individual diagnosis, a point developed at what an IQ test measures.
The commercial present is separate from that provenance. The current cognitron.co.uk site carries a copyright notice for H2 Cognitive Designs and describes over 100 validated paradigms tested by over 1,000,000 participants, with military, clinical and research institutions named as customers and institutional partnership as the primary model. Those are vendor stated figures. A buyer should not assume that the published Imperial College research automatically validates whatever configuration a commercial contract delivers today, and should ask which specific paradigms in the proposed battery carry published psychometrics.
8 Consumer Brain Training: Why It Appears in the Same Results
Brain training products rank for cognitive assessment queries because they contain assessment shaped screens, not because they are assessments. A training product measures you at the start so it can show you progress later, and progress is the thing being sold. That is a coherent product design, and it is structurally incompatible with the job of placing a person in a population.
Three properties make a training score unusable as an assessment. It is repeated by design, so practice effects are not noise to be controlled but the mechanism the product runs on. Its difficulty usually adapts to the user, so two people with the same displayed number may have answered entirely different item sets. Its reference group is the product's own user base, which is self selected in a way no vendor can correct after the fact.
The transfer question is separate and older. Improvement on a trained task is evidence about that task. Whether it generalizes to untrained tasks in the same domain, and whether it generalizes further to everyday functioning, are two additional claims that require their own evidence, and the strength of that evidence weakens sharply at each step. The honest position is that gains on the trained activity are easy to demonstrate and broad transfer is not. This site treats that question in full at can you improve your IQ.
None of this makes training products bad. A person who enjoys a daily reasoning puzzle is doing something harmless and possibly pleasant. The failure mode is administrative: an organization that adopts a training product because its dashboard looks like assessment reporting, then finds it has a set of engagement metrics and no defensible statement about anyone's ability. If the deliverable is a score someone will act on, the product has to come from one of the other three categories.
The reverse mistake also happens. Teams reject an entire category because they have seen a consumer product dressed in assessment language and concluded that all online cognitive measurement is that. The Mobile Toolbox psychometrics in the previous section are the answer to that objection: the difference between a marketing score and a measured one is whether the validation studies exist and can be read. For the consumer facing version of this distinction, see free versus validated IQ tests and are online IQ tests accurate.
9 Pearson Q-interactive and Q-global: The Supervised Route
Pearson operates two distinct platforms, and buyers routinely name the wrong one. Q-interactive is the digital delivery system for individually administered tests. The examiner works from one iPad while the examinee responds on a second, and the examiner is present throughout, controlling the session and recording responses. Pearson's own materials list more than 20 assessments in the library, including WAIS-5, WISC-V and WPPSI-IV for ability, KTEA-3, WIAT-4 and WRAT5 for achievement, CELF-5, GFTA-3, PPVT-5 and EVT-3 for speech and language, CVLT-3 and WMS-5 for memory, D-KEFS for executive function, and NEPSY-II and the RBANS Update for neuropsychology. Some tests still require physical materials such as blocks and response booklets, which Pearson explains is necessary to maintain construct equivalence with the paper versions. That library is a delivery system wrapped around instruments that exist independently of it, and the row by row inventory of intelligence batteries lists the same Wechsler scales beside the Stanford-Binet 5, the KABC-II and the DAS-II with every publisher, current edition, age range, administration mode and qualification level in the same fields, which is the comparison to settle before choosing software to deliver any of them.
Q-global is the web based platform for on screen administration, scoring and reporting, consolidating several legacy systems. It adds video proctoring so a clinician can conduct an assessment remotely while staying inside the normal workflow, described as supporting more than 40 assessments, along with sub account management, examinee records and searchable assessment histories.
The gate on both is qualification. Pearson assigns every professional assessment product a qualification level of A, B or C, with C the highest, and applies those levels to accounts based on existing credentials or a submitted qualification form. The Wechsler scales sit at the restricted end of that structure. This is the single most important structural fact about the category: the platform is not the constraint, the examiner is. Full detail on the levels and who satisfies them lives at the qualification level breakdown.
The commercial model reflects the institutional buyer. Q-interactive combines an annual per user license, varying by test bundle, with per subtest usage charged either pay as you go or prepaid at volume. Q-global offers one, three or five year subscriptions for selected products alongside pay per report pricing that varies by service type. Neither is a self serve research tool, and neither is priced to be one.
When this category is not optionalIf the result will be read by a school placement committee, a court, a disability determination or a clinician making a diagnosis, the supervised route with a qualified examiner is the requirement, and no unsupervised platform substitutes for it. Background on the current Wechsler scale is at what is the WAIS-5, and on the leading nonverbal alternative at what is Raven's 2.
10 ACIS Professional: A Normed Battery You Administer Remotely
ACIS Professional sits in the fourth category, the normed ability battery, on the unsupervised side of it. It administers the same instrument that individual examinees take, to participants you invite. The structure is 20 subtests across six primary cognitive domains organized on the Cattell Horn Carroll model: comprehension knowledge (Gc), fluid reasoning (Gf), quantitative reasoning (Gq), visual spatial ability (Gv), working memory (Gwm) and processing speed (Gs), reported as VCI, FRI, QRI, VSI, WMI and PSI. Subtest scaled scores use a mean of 10 and a standard deviation of 3, composites a mean of 100 and a standard deviation of 15. The domain structure is described at the CHC model and the score metric at standard deviation 15 explained.
The published psychometrics come from a technical analysis set of 2,750 complete records, interpreted within an adult reference frame built from 3,243 English speaking records covering ages 16 to 90. The Full Scale IQ composite has an omega of .9886 and a g loading of .958, with a standard error of measurement of about 1.60 IQ points. The higher order g confirmatory model reports CFI .9761, TLI .9726, RMSEA .0406, SRMR .0217 and a chi-square of 916.703 on 166 degrees of freedom. Every one of those figures carries the same three qualifications: the sample is self selected rather than census based, administration is unsupervised, and the reference frame is a modelled adult frame rather than a census sample. The full documentation is at the technical manual.
Operationally, a free Professional account includes three Quick administrations, after which administrations are paid at the same prices individual examinees pay. You create administrations in a workspace, share private participant links, choose whether each participant sees their own results, and export verified scores as CSV. Credits are charged when links are created rather than when a participant starts, and an unopened administration can be revoked. Participants exist in the system as pseudonymous codes, not as people, so no participant personal data is collected.
20 subtests
Across six CHC domains, yielding six primary indices plus a Full Scale composite.
3 free
Quick administrations included with a free Professional account, no credits required.
No participant PII
Participants appear as pseudonymous codes, and scores export as CSV.
Boundaries, stated plainlyACIS is self administered and unsupervised. It is not a clinical instrument, it is not proctored, and it does not replace an individually administered evaluation by a licensed professional. It is not for diagnosis, hiring decisions, accommodation determinations or high IQ society admission. It suits research, screening and educational contexts where an unsupervised, normed, multi domain profile is what the question needs.
This table is sorted by category, because sorting by score would reproduce exactly the error this page exists to correct. Read down to your category first, then across.
Platform
Category
What it outputs
Who administers
Question it answers best
Creyos
Clinical screener
Domain scores plus behavioral health questionnaires, structured for the health record
Clinician or clinic staff
Has this patient changed, and can I document and bill it
CANTAB
Clinical trial and research battery
Sensitive per task measures across four domains
Trained study or clinical staff
Did this arm differ from that arm
Cognitron
Large scale remote research battery
Performance scores across many paradigms, modelled at group level
Research team, participant self administers
How does a very large sample distribute and differ
NIH Mobile Toolbox
Fixed normed research battery
Per measure scores with published reliability and validity
Research team, participant self administers on a phone
How does my cohort compare on established measures
Gorilla
Research task builder
Trial level data from tasks you design
You, as the experimenter
Did my manipulation work
Testable
Research task builder with participant pool
Trial level data plus recruited, identity verified participants
You, as the experimenter
Did my manipulation work, and where do I find people
PsyToolkit
Research task builder, free
Trial level data from a paradigm and scale library
You, or your students
Can I run and teach this study at no cost
Pearson Q-interactive
Supervised publisher battery
Standardized scores from WAIS-5, WISC-V and other published tests
A qualified examiner, present in the session
What is this individual's standing, defensibly
Pearson Q-global
Publisher scoring and reporting platform
Scored reports, on screen administration, video proctoring
A qualified professional
How do I score, report and store results properly
ACIS Professional
Normed ability battery, unsupervised
Six index profile plus a Full Scale composite in the IQ metric, CSV export
You invite, the participant self administers remotely
Where do my participants sit on an adult reference frame
Notice how few of these rows contain the same words in the output column. That is the whole argument. When a listicle ranks these ten against one another with a single numeric score, it is averaging across a column whose entries are not commensurable, and the resulting order tells you about the rubric's assumptions rather than about the products.
Notice also that only two rows produce a composite in the IQ metric, and they sit at opposite ends of the administration spectrum. That is not a coincidence. A composite is worth reporting only when a reference frame makes it interpretable, and building a reference frame is the expensive commitment most platforms in the other categories deliberately decline to make. The consumer side of that same distinction is mapped at the best online IQ tests and the best IQ test, both written for someone testing themselves rather than administering to others.
12 What to Verify Before You Sign, in Any Category
Once the category is settled, the remaining diligence is the same everywhere, and most of it is answerable from public documents before you ever speak to a salesperson. The five questions below are ordered by how often the wrong answer kills a project after the money is spent.
Question to ask
Why it decides the purchase
Where a good answer lives
What reference sample interprets a score, and how was it recruited?
Without it, an individual number means nothing and only group comparisons are possible
A technical manual or a published norming paper, not a marketing page
What reliability is reported, for which specific score, by which method?
Internal consistency, test retest and conditional error answer different questions and are not interchangeable
Per measure tables, as in the Mobile Toolbox papers cited above
What validity evidence links this measure to a named external instrument?
A moderate published correlation is informative; an unquantified claim of validation is not
Peer reviewed convergent and discriminant coefficients
Can I export raw, trial level data, and who owns it?
Decides whether reanalysis, replication and secondary use are possible at all
The contract and the data processing terms, not the feature list
What happens under unsupervised administration?
Effects attenuate remotely, as the Gorilla flanker comparison showed directly
The vendor's own remote validation, if it exists
Four buying mistakes account for most wasted spend. Buying across categories, which this page has covered at length. Buying a builder and expecting norms, which produces raw counts nobody can interpret. Buying a composite when the hypothesis was about one domain, which wastes administration time on subtests that do not bear on the question. And buying a domain profile when the hypothesis was about general ability, which leaves you unable to make the claim you set out to make.
Two questions are deliberately not answered here because sibling pages own them. What each platform costs, including how per participant, per seat and subscription models compare once you model realistic volume, is at the cost of cognitive testing platforms. Who is permitted to administer which instrument, and what the A, B and C qualification levels actually require, is at who can administer an IQ test. Treat both as prerequisites rather than as follow up reading.
13 Judging Any Platform Against the Published Standards
Every question in this article is a restatement of one principle from the professional standards: validity belongs to an interpretation for a specific use, not to a test or a platform in the abstract. The Standards for Educational and Psychological Testing, published in 2014 by the American Educational Research Association, the American Psychological Association and the National Council on Measurement in Education, is explicit that a test author must specify the intended interpretations and uses, and that evidence must be assembled for those uses rather than asserted generally. That single sentence dissolves the ranking listicle problem. A platform cannot be better than another platform. An interpretation can be better supported for a purpose.
The same framework supplies the checklist this page has been working through. Evidence based on test content asks whether the tasks sample the domain claimed. Evidence based on internal structure asks whether the reported composites match how responses actually cohere, which is what a confirmatory factor model reports. Evidence based on relations to other variables asks whether scores correlate with the instruments and outcomes they should, and diverge from those they should not. Reliability is required for each reported score, not for the instrument as a whole. Fairness runs through all of it, as a requirement rather than an appendix.
Two further documents matter for anyone administering remotely. The American Psychological Association's testing standards resources set out the professional obligations attached to test use, including the responsibility of the user to work within the limits of their qualification. The International Test Commission guidelines address computer based and internet delivered testing directly, including the control levels of administration and what an open, unsupervised session can and cannot support.
Applied to this page's ten platforms, the standards produce a simple discipline. State the interpretation you intend to make. Identify which category can support it. Ask the vendor for the evidence that specific interpretation requires. Accept the limitations the evidence carries, including the ones this site states about its own instrument: unsupervised administration, a self selected normative sample, a modelled adult reference frame, and no diagnostic or high stakes use. A platform that states its boundaries has told you where its evidence ends. A platform that states none has told you nothing at all.
The closing testBefore signing, write the sentence you intend to say when the data arrives, then ask whether the Standards would call that sentence a supported interpretation for this use. If the honest answer is no, the problem is the category, and no amount of vendor comparison inside the wrong category will fix it.
14 Frequently Asked Questions
What is a cognitive assessment platform?
It is software used to administer cognitive tests to other people and collect the results centrally. The term is not standardized, so it currently covers clinical screeners, research platforms, consumer training products and normed ability batteries.
Why do ranking lists of cognitive assessment tools disagree so much?
Because they score products from different categories on one rubric. Change the weighting slightly and the order flips, since the products were never competing for the same job in the first place.
Which cognitive assessment platforms give an IQ score?
Very few. Pearson Q-interactive delivers published Wechsler scales through a qualified examiner, and ACIS Professional delivers a normed composite remotely. Most research and clinical platforms report per measure or per domain results instead.
Is Creyos the same thing as Cambridge Brain Sciences?
Creyos is the current brand for that lineage of twelve cognitive tasks, and its scientific leadership traces to Adrian Owen's laboratory work. The product today is positioned as a healthcare platform with record integration rather than as a research tool.
Can I use a research task builder to screen job applicants?
No. A builder gives you raw performance data with no reference sample, so there is no defensible way to say an applicant scored above or below any standard. Employment testing also carries legal requirements that no builder addresses.
What is the difference between a task builder and a fixed research battery?
A builder is an authoring tool where you design the paradigm yourself and own the resulting design. A fixed battery is a set of published instruments you administer as specified, with per measure psychometrics already available.
Does the NIH Mobile Toolbox cost anything to use?
The platform does not publish a simple consumer price on its main page, and access terms depend on the study arrangement. Budget questions across all these platforms are covered on the dedicated cost page rather than here.
Is CANTAB better than the Mobile Toolbox?
They serve different buyers. CANTAB is engineered for regulated trials and clinical programs where sensitivity to change is the endpoint. The Mobile Toolbox is engineered for remote, self administered research at population scale.
Why do cognitive research platforms avoid reporting a single composite?
Many were built on models that treat cognition as several separable systems rather than one. Reporting a composite would also require a normative reference frame, which is a large and expensive commitment most research platforms decline.
Can Gorilla or PsyToolkit produce a valid intelligence measure?
They can run the tasks, but validity comes from norms, structural evidence and reliability work that you would have to perform yourself. Building the tasks is the small part of that project.
What does Q-interactive require that Q-global does not?
Q-interactive is built around an examiner physically present with the examinee, working from a second device throughout the session. Q-global focuses on screen administration, scoring and reporting, and offers video proctoring for a subset of assessments.
Do I need a qualification to buy any of these platforms?
It depends entirely on the instrument, not on the software. Published clinical tests carry qualification levels that gate access, while research builders and most remote batteries do not impose the same restriction.
Is unsupervised remote testing acceptable for research?
It is widely used and widely published, but it changes what the data can support. Effects measured remotely tend to be attenuated relative to laboratory administration, so sample sizes must be planned against the smaller expected effect.
How large should a normative sample be before I trust a platform's scores?
Size matters less than composition and documentation. A modest sample that is described honestly, with stated recruitment and exclusions, supports better interpretation than a very large sample of unknown origin.
What is the most common expensive mistake in this market?
Choosing a vendor before choosing a category. Every downstream comparison is then made on features that do not bear on whether the product can answer the question that justified the purchase.
Can a clinical screener replace a full neuropsychological evaluation?
No. Screeners are built to flag change and support documentation within a care pathway. A full evaluation involves history, observation, multiple instruments and clinical judgment that no platform supplies.
Do these platforms let me export raw trial level data?
Research builders and research batteries generally do, since reanalysis is their purpose. Clinical platforms often export at the report level instead, so confirm the export format and data ownership terms before signing.
What does ACIS Professional not do?
It does not diagnose, it does not proctor, and it does not produce a result intended for hiring, accommodation determinations or high IQ society admission. It provides a normed multi domain profile for research, screening and educational use.
How many administrations does a free ACIS Professional account include?
Three Quick administrations, provided as courtesy rather than as credits. Unopened administrations can be revoked, and a revoked courtesy administration returns to the allowance.
Should I compare platforms on price first?
Only after the category is fixed. Comparing prices across categories ranks products that cannot substitute for one another, which is how teams end up owning something cheap that answers nothing.
Where do consumer IQ test rankings fit in this comparison?
They serve a different reader entirely, someone testing themselves rather than administering to others. The consumer ranking pages on this site cover that audience and do not overlap with the platform decision described here.
Take the assessment
You get a profile, not a number
ACIS measures six CHC domains across 20 subtests and reports each one with its own normed score and confidence interval, so you can see where you are strong and where you are not.