International IQ Test Review: What 16M Results Do Not Prove
International IQ Test publishes more methodology than many consumer quizzes and offers a free estimate in more than 40 languages. Its matrix focused design and self selected online database still create important limits for personal and country level conclusions.
International IQ Test evaluated separately from the country rankings produced from its online participant database.
0 The Short Answer
The International IQ Test publishes far more about its own methods than almost any other free online IQ test, and nearly all of that evidence was produced by the company that sells the test. Both halves of that sentence matter. Most consumer IQ sites publish nothing beyond a marketing page, so a provider that posts a standardization report with sample sizes, year by year distributions and a documented weighting procedure has done something genuinely unusual and genuinely useful. That work deserves credit before any criticism arrives.
The International IQ Test is a 40 item, Raven inspired matrix test that takes roughly 20 to 30 minutes, runs in more than 40 languages, and shows your score at the end at no charge. Its scoring model is calibrated to a mean near 100 with a standard deviation near 15. Its highest possible score is 142. Those are the product facts, and they are stated plainly on the provider's own pages rather than buried.
The part that needs careful reading is the headline figure of roughly 0.943, described on the site as the association between the test's score and a general intelligence factor. That number is real, the report explaining it is public, and the report itself says exactly what the number is: an internal analysis in which a general factor is extracted from the same 40 answers the score is built from. It is a check that the scoring algorithm faithfully summarizes its own items. It is not a comparison against any external measure of cognitive ability, and the provider says so in its own transparency box. Understanding why those two things are different is worth more to you than any verdict on this particular product, so most of this page is spent on it.
Who wrote this and why that matters
This review was written by the team behind ACIS, which sells a competing cognitive assessment. Read it as an informed competitor's evaluation rather than a neutral third party audit. Everything below is checkable: we describe only what the provider publishes publicly, we date every price and every figure, and we link to the source pages so you can confirm each claim yourself. Where we could not find something, we say we could not find it rather than assuming it does not exist.
1 What the Product Actually Is
Strip away the framing and the International IQ Test is a single, timed, nonverbal reasoning task delivered in the browser. The published test page lays out an incomplete visual pattern with six lettered answer options, and the instruction is to choose the missing figure. The methodology page states that the design is inspired by Raven's Progressive Matrices, the format John C. Raven introduced in 1936, chosen because a pattern completion task can be presented identically to a Vietnamese speaker and a Norwegian speaker without translating a single item stem.
The administration itself is short and unfussy. A running clock and a progress counter sit above the item on the published page, the counter reads out of 40, and the provider's own guidance says the whole thing usually takes 20 to 30 minutes and that timing has only a small effect on the result. There is no registration wall before the test. At the end, an optional form asks for a display name, an email address, country, gender, year of birth, field of study and education level, and the page explains what each field is used for: the email deduplicates results and lets you retrieve your score later, and the demographic fields feed the comparison statistics and the annual rankings.
Language coverage is the standout practical feature. The language switcher offers roughly 45 options, from Icelandic and Latvian to Tagalog, Farsi, Hebrew and both written forms of Chinese. Very few cognitive instruments of any kind, free or paid, are reachable in that many languages, and for a nonverbal task the translation burden is limited to instructions rather than items. If your requirement is simply to give the same reasoning task to people across a dozen countries this week, there are not many alternatives with this reach.
What you get back is a score plus positioning statistics: where the number falls against the general population and against your own age group. The provider describes this as the free result, and the score appears immediately without payment.
2 What the Provider Publishes, in Numbers
Four documents carry the technical claims, all last updated in late December 2025: a methodology and mission statement, a reliability summary, a standardization and calibration report, and an analysis of the score's relationship to a general factor. Reading all four takes about fifteen minutes, and doing so puts you ahead of almost anyone else discussing this test online. Here is what they report.
40 items
The g factor analysis describes converting each participant's answers into 40 binary variables, which fixes the test length precisely.
66,032 per year
The standardization report's effective annual sample after country weighting, anti bot screening and deduplication, against a target of 80,000.
Mean 99.75 to 100.86
Observed means across the 2020, 2021 and 2022 weighted samples, with standard deviations of 15.12, 15.15 and 15.49.
50,000 per year
Sample size in each of the two annual g factor analyses, drawn from 2024 and 2025 administrations with one unique email per participant.
0.9437 and 0.9429
The internal g loadings reported for 2024 and 2025, with confidence intervals so tight they span roughly one thousandth.
16,965,148 results
The database total from January 2018 to December 2025 after annual deduplication by both email and IP, from 21,174,884 raw results.
Two features of this disclosure deserve explicit credit. First, the provider publishes three different counts of its own database rather than one flattering one: 21.17 million raw results, 19.70 million after email deduplication, and 16.97 million after email and IP deduplication. Providers who want to look big quote the raw number and stop. Second, the reports state what they do not establish. The standardization report says it does not cover cognitive validation or external clinical validation. The g factor report says it does not replace clinically supervised comparison against instruments such as the Wechsler scales. Those sentences were written by the seller, and they are the correct sentences.
3 How to Read the 0.943 Figure Correctly
This is the number that travels. It appears on the homepage, on the methodology page and in third party write ups, usually compressed into something like "correlates 0.94 with general intelligence," and in that compressed form it is badly misleading. The full report is clear about what was actually computed, so the misreading is the internet's fault rather than the provider's, but it is worth taking apart slowly because the same trick appears across the industry.
Here is the procedure as the report describes it. Each participant's 40 answers are coded 0 for incorrect and 1 for correct. A general factor score is then estimated for each participant by extracting the first principal component from that 40 item response matrix. Finally, that principal component score is correlated with the IQ score the site's own algorithm produced from the same 40 answers. The result is roughly 0.943 in both 2024 and 2025.
Read that sequence again and notice that only one set of data ever enters it. The general factor and the score being validated are two different summaries of the same 40 responses from the same person on the same afternoon. A high correlation between them tells you that the scoring algorithm is a faithful, roughly linear compression of its own items, which is a real and worthwhile thing to verify. It does not tell you that the test agrees with any other measure of reasoning, because no other measure was involved at any point.
The report contains its own tell, and to the provider's credit it prints it: the correlation between that principal component and the plain raw total score is 0.9874, higher than the correlation with the algorithm's output. In other words, simply counting correct answers tracks the extracted factor even more closely than the proprietary scoring does. That is exactly what you would expect from a within instrument consistency check, and it shows that the 0.943 figure is not measuring the sophistication of the scoring model. Almost any sensible summary of 40 correlated binary items would land in the same neighborhood.
One more line from the same report keeps the arithmetic honest: the first principal component accounts for about 15 percent of the variance in the binary items. That is an unremarkable figure for a set of 40 dichotomous items and it is not in tension with the 0.943, because the two quantities answer different questions. But it does mean that anyone reading 0.943 as "94 percent accurate" has misunderstood the statistic by a wide margin. Accuracy against what, measured how, is the question the number does not answer.
4 Why Provider Run Validation Needs Outside Replication
Nothing above accuses anyone of anything. The analysis is disclosed, the method is described in enough detail to be criticized, and the limits are stated on the same page. That is better practice than the great majority of the market. The question this section answers is narrower and more useful: why does the same finding carry less weight when the seller produces it, and what would raise it?
Start with the structural reason, which has nothing to do with honesty. Any analysis involves choices: which years to sample, which filters to apply, how to define the comparison factor, which of several defensible estimates to feature. When a result comes out flattering, nobody goes looking for the choice that made it so. When it comes out badly, the analysis gets rerun. Neither step requires bad faith, and the aggregate effect across an industry is that self reported figures skew high. This is why medicine separates the sponsor from the trial statistician, and why psychometrics treats a coefficient replicated by an outside lab as a different class of evidence from the same coefficient reported in house.
Then there is the selection reason. A provider decides which analyses to publish. We can see two reports here and they are informative. We cannot see whether other analyses were run and shelved, and that is not a suspicion about this company specifically; it is true of every organization that publishes its own numbers, including ours.
So what would independent replication look like for a test like this one? The bar is concrete and not especially high:
An external criterion. A sample of people who take the online test and also sit a supervised battery, with the correlation between the two published alongside the sample size and recruitment method. Concurrent evidence against an outside instrument is the thing a within instrument analysis cannot substitute for.
Investigators without a stake. Analysis run or at least audited by people who do not benefit from the answer, ideally with the analysis plan registered before the data are seen.
Test retest data. The same people, the same instrument, a documented interval, and the resulting correlation. This is the coefficient that tells an individual how much their own number would move if they took the test again next month.
Reanalyzable material. Item level data, or at minimum an item difficulty and discrimination table, so that someone else can reproduce the reported figures rather than take them on trust.
Apply that standard evenly and it cuts in every direction, ACIS included. Our own accuracy documentation is also produced in house, which is a limitation on our claims exactly as it is a limitation on this provider's. The useful reader habit is not "distrust sellers." It is to ask, of any coefficient you meet, who computed it, against what external thing, and whether anyone outside the building has reproduced it.
5 The Norm Sample and What Country Weighting Does
The standardization report is the strongest document on the site, and it repays a close look. The procedure was to build three separate annual samples, one each from 2020, 2021 and 2022, in which every country contributes results in proportion to its share of world population. China represents about 18.89 percent of the world's people, so a target sample of 80,000 reserves 15,112 slots for China. Anti bot screening and duplicate filtering run first. The resulting distributions came out at mean 100.86 with a standard deviation of 15.12 in 2020, mean 99.75 with 15.15 in 2021, and mean 99.82 with 15.49 in 2022.
That is a well specified and reproducible design, and the stability across three years is meaningful. It is also a much weaker claim than "normed on the world," and the report is candid about why. Only 82.54 percent of the target representation could be filled, giving 66,032 of the intended 80,000 per year, because countries in the remaining 17.46 percent had too few administrations to meet their quota and were dropped. Those excluded countries are not random. They are disproportionately places with lower connectivity, which is precisely where a globally representative sample would most change the answer.
The deeper issue is the one country weighting cannot touch. Quota matching fixes how many results come from each country. It does nothing about who inside each country chose to click a free online IQ test and finish 40 timed matrix items. Volunteers for cognitive tests differ systematically from their neighbors in education, curiosity, screen access and confidence, and no amount of geographic reweighting corrects a selection effect that operates within every cell of the table at once. The provider names the related limitation, noting that the test reaches internet users, about 74 percent of the world population in 2025 by the ITU figure it cites.
The practical consequence for you as a test taker is modest but real. Your score is a position within a large volunteer population that has been geographically reshaped to resemble the world, not a position within the world. If you want to understand how much interpretive weight a percentile can bear, our page on reading IQ percentiles works through the same arithmetic in more detail.
6 What 16 Million Results Buys, and What It Does Not
Size is the most quoted fact about this product and the most misread. A database of nearly 17 million deduplicated results is a genuine asset, and it buys three specific things. It makes item level statistics extremely stable, so difficulty and discrimination estimates for those 40 items are known with a precision most test publishers would envy. It makes rare score regions populated rather than extrapolated. And it makes annual comparisons possible, which is why the provider can publish year over year rankings by country, age, education level and field of study at all.
What size does not buy is representativeness. This confusion is the single most common error in consumer psychometrics, so it is worth stating flatly: a sample can be enormous and still be wrong about the population, because the error introduced by who volunteers does not shrink as the sample grows. It stays exactly where it is while the confidence interval around it collapses, which produces the worst possible combination, a biased estimate reported with great precision. The classic demonstration is the 1936 Literary Digest poll, which collected 2.4 million responses and called the American presidential election for the wrong candidate, while George Gallup got it right with a few thousand carefully selected ones.
Size also does not buy construct coverage. Seventeen million administrations of a matrix test produce a very well characterized matrix test. They produce nothing at all about vocabulary, working memory or processing speed, because those abilities were never sampled. Precision in one place cannot be spent somewhere else.
None of this makes the database uninteresting. The country and age tables are one of the few large, publicly described cross national datasets on any cognitive task, and treated as what they are, a description of who takes this test where, they are a legitimate resource. The failure mode is treating them as national IQ estimates, which the sampling cannot support, and which our page on average IQ figures and where they come from unpacks separately.
7 One Ability, Measured Well, Is Still One Ability
Matrix reasoning is a strong indicator of fluid reasoning, the ability to solve novel problems without relying on stored knowledge. It is a legitimate and well studied task type, and a 40 item matrix set administered under time pressure is a reasonable way to sample it. The provider states the limitation directly on its methodology page: Raven style matrices do not measure verbal, social or emotional dimensions.
The evidence on how far a matrix score can be stretched is worth knowing precisely. A 2015 analysis by Gilles Gignac in the journal Intelligence, titled "Raven's is not a pure measure of general intelligence," fitted bifactor models across several large samples and found that Raven's shared roughly 50 percent of its variance with a general factor, about 10 percent with a fluid reasoning group factor distinct from that general factor, and carried around 25 percent test specific reliable variance. Gignac's recommendation was explicit: researchers should not rely on Raven's alone when a valid estimate of general ability is the goal, and he offered a four subtest Wechsler combination with a general factor validity coefficient of .93 in about 14 minutes as one alternative.
Hold that next to the 0.943 discussed earlier and the difference becomes vivid rather than contradictory. Gignac estimated how much a matrix test shares with a general factor built from a diverse battery of different tasks. The provider's report estimated how much its score shares with a factor extracted from its own matrix items. The first quantity is the one that answers "does this predict general ability," and the external literature puts it far below 0.94 for matrix tests as a family. The second quantity is a coherence check. Same style of statistic, different universe.
For a test taker the implication is practical. A single strong number here means your fluid reasoning performed well on this task on this day. It says nothing about your vocabulary and general knowledge, nothing about how much you can hold and manipulate in working memory, nothing about how fast you execute simple decisions under time pressure, and nothing about quantitative reasoning. Those domains dissociate in real people all the time, which is the entire reason full batteries report profiles rather than a single figure. If the distinction between task specific strength and general ability is new to you, our explainer on fluid and crystallized reasoning is the shortest route into it.
8 The Ceiling at 142 and Who It Excludes
The provider states that the highest possible score on the International IQ Test is 142. Publishing a ceiling at all is a mark of seriousness, since most consumer tests quietly let their scoring curve run off into fantasy territory and hand out 150s to flatter their users. Naming the maximum tells you where the instrument stops resolving differences, which is information you cannot get from most competitors at any price.
It also tells you who this test cannot serve. On the standard scale with a mean of 100 and a standard deviation of 15, here is where the ceiling sits.
142 is 2.8 SD up
Roughly the top 0.26 percent of the scale, about 1 person in 390. Everything above that compresses into one score at the top of the range.
145 is unreachable
Three standard deviations, about 1 in 740, is outside what this instrument can report at all, whatever the taker's actual ability.
160 is far outside
Four standard deviations, roughly 1 in 31,000. A test with a 142 maximum has no mechanism for distinguishing anyone in this region.
A ceiling this low is a design choice rather than a defect. Resolving scores above roughly 2.5 standard deviations requires items difficult enough that most takers fail them, which lengthens the test and worsens the experience for the 99 percent of users who will never approach that range. Capping at 142 keeps the instrument short and pleasant for the population it actually serves. The cost falls entirely on high scoring takers, who will hit the ceiling and learn only that they are somewhere above it. If your reason for testing is to characterize performance in the top fraction of a percent, this is the wrong tool and no amount of retaking will change that. Our comparison of what different IQ tests are built to do covers which instruments carry usable range at the top.
9 Price, Add-Ons and What You Actually Receive
As of August 2026, the front page of the International IQ Test states that you will get your IQ score for free with no payment required, and the Terms of Use and Privacy Policy page repeats it: the test and the basic results, including the IQ score, are available free of charge. The site's own description of the free result is a score plus comparison statistics against the general population and against your age group. In a category where the standard pattern is 40 timed minutes followed by a paywall demanding twenty or thirty dollars to see the number you just earned, delivering the score at no cost is the single most consumer friendly thing about this product.
There is a paid layer, and here the disclosure is thinner. The same Terms page names FastSpring as the merchant of record for optional add-ons and lists a Santa Barbara, California mailing address, but it does not say what the add-ons are or what they cost, and we could not locate any public page on the site as of August 2026 that publishes a price list or describes the paid product. An archived version of that same Terms page from 21 June 2024, viewable in the Internet Archive, was more specific: it described paying to validate your result through the FastSpring reseller so that you would learn which answers were right and wrong. Whether that is still the offer today we cannot confirm from the public pages, and prices and packages change, so the checkout you see at the moment of purchase is the only authority on what you will be charged.
Two practical notes follow. First, if all you want is a number and a percentile, you should not need to pay anything, and you should be suspicious of any flow that suggests otherwise. Second, because the price of the optional layer is not published in advance, budget for the possibility that the item level breakdown costs something before you decide it is essential. For context on what other providers charge for comparable outputs, our breakdown of what IQ testing costs lists current prices across the range from free web tests to supervised clinical assessment.
10 Data Handling and Retest Integrity
The privacy section is short, specific and better than the category norm. It states that the only directly identifying data collected are your email address and your IP address, that the email exists to let you retrieve your result and to deduplicate and detect bots, and that the remaining fields you may supply, display name, gender, year of birth, field of study, education level and country, are used to generate the statistics shown with your result and the aggregate displays on the site. It states that stored identifying data is not sold or shared and that no marketing email is sent. It offers deletion under the GDPR by writing from the address used to create the result. It also discloses that Google and its partners may serve advertising with cookies on some pages, which is the honest way to describe an ad supported model.
Providing the email is optional, which is worth knowing before you decide. Skipping it costs you the retrieval feature and the age group comparison, and it keeps your result out of the annual statistics.
On integrity, the provider applies deduplication by email and IP for its published statistics and says stricter internal anti cheating filters exist but are deliberately not described, so that people cannot adapt to them. That reasoning is sound for protecting a public dataset. It is worth being clear about what those controls do and do not protect. They protect the aggregate rankings from bot traffic and mass retaking. They cannot protect your individual result from the ordinary realities of unsupervised testing: a second monitor, a friend in the room, an interruption, a bad night's sleep, or simply having seen a matrix test before.
That last one has a measurable size in the literature. A 2007 meta-analysis by Hausknecht and colleagues in the Journal of Applied Psychology found consistent score gains from retesting on cognitive ability measures, with practice and coaching both contributing. The provider's own advice matches the evidence: wait at least a year before retaking, and treat your first score as the most trustworthy one. Following that advice is free and improves your data more than anything else you can do.
11 Strengths and Limitations Side by Side
No score out of fifty appears on this page, because a number invented by a competitor is worth nothing to you unless every criterion behind it is written out and checkable. What follows instead is the evidence laid out in both directions, each row traceable to something the provider publishes or something we could not find published.
Dimension
What supports it
What limits it
Published methodology
Four dated technical pages covering mission, reliability, standardization and the general factor analysis, with samples and procedures described.
All of it is produced in house. No outside group has reproduced any figure, and no analysis plan appears to be registered in advance.
Score calibration
Country weighted samples across 2020, 2021 and 2022 land within a point of a mean of 100 with standard deviations near 15, three years running.
Coverage reached 82.54 percent of the target, and quota matching cannot correct who volunteers inside each country.
Evidence of internal structure
A g loading near 0.943 replicated across two annual samples of 50,000, with confidence intervals reported.
The comparison factor is extracted from the same 40 items, so it demonstrates coherence rather than agreement with an outside instrument.
Reliability in the usual sense
The provider defines and defends scale stability across years, which is a real form of consistency.
We could not locate a published internal consistency coefficient, a standard error of measurement or a test retest correlation on any public page.
Cost and access
Score delivered free at the end, no registration required to start, roughly 45 language versions.
The optional paid layer is named in the terms through its merchant of record but its contents and price are not published on the site.
Scope of the result
Fluid reasoning is sampled with 40 items and a stated ceiling of 142, and the limits are disclosed rather than hidden.
Verbal comprehension, working memory, processing speed and quantitative reasoning are never sampled, and the ceiling excludes the top fraction of a percent.
12 Who It Fits, and Where ACIS Sits
Take the International IQ Test if you want a free, fast, language flexible estimate of nonverbal reasoning and you are comfortable reading the result as exactly that. It is a reasonable choice for casual curiosity, for a rough sense of where you land on pattern reasoning, for comparing your own result across a long interval, and for anyone who needs a task that works the same way in Vietnamese and Finnish. Within that brief it does its job and charges nothing. Readers weighing alternatives can set this profile against our findings on IQ Test Academy and IQTest.com, two products that publish far less than this one does.
Look elsewhere in three situations. If a decision hangs on the outcome, an educational placement, a workplace question, a clinical concern, only a supervised assessment by a qualified professional carries the weight, and every online instrument including ours says so. If you need a profile across domains rather than one number, a single ability test structurally cannot supply it. And if you expect to score near the top of the distribution, the 142 ceiling will cut you off.
ACIS is our own product and we will keep the description short and checkable. It is an online, self administered battery of 20 subtests across six CHC domains, reporting verbal comprehension, fluid reasoning, visual spatial processing, working memory, processing speed and quantitative reasoning as separate indices with confidence intervals, alongside a composite. It is not a clinical instrument, it does not diagnose anything, no institution is obliged to accept its output, and its norms are documented in a public technical manual that we wrote, which places our evidence under the same replication caveat this page has applied throughout. The honest comparison is not that one product is better. It is that a matrix test answers one question cheaply and a broad battery answers a different, wider question at a different cost, and knowing which question you are asking settles the choice. Our side by side of supervised and online assessment covers the third path.
Yes for the score itself. The provider's terms state that the test and basic results are available at no charge, and the score appears at the end without payment. Optional add-ons are handled by a third party merchant of record.
How many questions does it have?
Forty. The provider's own factor analysis describes coding each participant's answers as 40 binary variables, and the item counter in the interface runs to 40.
How long does it take?
The site says 20 to 30 minutes for most people and notes that timing has only a small influence on the outcome. A clock runs on screen throughout.
What is the highest score you can get?
142. The provider names this maximum openly, which means performance above roughly 2.8 standard deviations cannot be distinguished by this instrument.
Does a 0.943 correlation mean it is 94 percent accurate?
No. Correlation coefficients are not accuracy percentages, and this particular coefficient compares two summaries of the same 40 answers rather than comparing the test against anything external.
Has any outside group verified the reported figures?
We could not find one. Every technical analysis we located on the site was authored by the provider, and we found no peer reviewed paper or third party dataset reproducing the numbers.
Does the provider publish a reliability coefficient?
Not in the conventional sense. The reliability page argues from scale stability and internal structure rather than reporting an internal consistency statistic or a standard error of measurement.
Is a Raven style test a full IQ test?
No, and the provider agrees. Matrix items sample fluid reasoning. Verbal knowledge, memory span, decision speed and numerical reasoning are outside what the format can reach.
Why does the academic literature give matrix tests lower general factor loadings?
Because those studies estimate the shared factor across several different task types, while an internal analysis extracts it from one task type. The two procedures are answering different questions.
Are the country rankings a measure of national intelligence?
They describe the people from each country who chose to take this specific test online. That is a self selected group, and no reweighting fixes who volunteers within a country.
Does a bigger database make the norms better?
It makes them more precise, not less biased. Volume shrinks the margin of error around an estimate without moving the estimate itself toward the truth.
Can I take it without giving an email address?
Yes. The form at the end is described as optional. Skipping it means you cannot retrieve the result later and your data will not feed the comparison statistics.
What does the provider do with the email?
Its policy says the address is used for result retrieval, deduplication and bot detection, that identifying data is not sold or shared, and that no promotional mail is sent.
Can I delete my result?
The policy offers deletion under GDPR if you write from the same address used to create the result. Deletion requests are handled through the contact address published on the site.
How soon can I retake it?
The provider recommends at least a year and says your first attempt is usually the most trustworthy. Research on retesting supports that caution, since familiarity alone lifts scores.
Will this result count for Mensa or any admissions process?
No. Unsupervised web tests are not accepted as qualifying evidence by high IQ societies, schools or employers. Each organization names the instruments it accepts.
Is it suitable for children?
The norms are described as calibrated to the global population rather than age banded for children, and any question about a child's development belongs with a qualified professional.
What if my score seems too low?
Single administrations move around for ordinary reasons: fatigue, interruptions, unfamiliar format, a slow start against the clock. One session on one task type is a data point, not a summary of you.
How does it compare with a supervised assessment?
A clinician controls the environment, verifies identity, observes how you work and combines several instruments with background information. None of that is possible in a browser tab.
Does ACIS apply the same scrutiny to itself?
It should, and we say so on this page. Our figures also come from our own analyses, our manual is written by us, and outside replication would strengthen our claims exactly as it would strengthen this provider's.
What single habit improves how you read any test's claims?
Ask what the number was compared against. A coefficient computed inside one instrument, a coefficient against a supervised battery and a coefficient against a real world outcome are three different grades of evidence wearing the same notation.
14 Sources Behind This Page
A review is only as good as the yardstick behind it. These sources document the measurement standards this review applies and the consumer practices it checks for.
American Psychological Association. The Standards for Educational and Psychological Testing, the joint AERA, APA and NCME framework that legitimate tests are built and evaluated against.
Buros Center for Testing. The independent center whose Mental Measurements Yearbook has reviewed commercial tests since 1938. The reference point for claims about test quality.
American Psychological Association. Understanding psychological testing and assessment. What professionally administered testing involves and how it differs from informal quizzes.
Federal Trade Commission (2022). Bringing Dark Patterns to Light. The FTC staff report on design practices that obscure subscriptions and charges, worth checking any paid online test against.
Pearson Clinical Assessment Scientific Council (2023). Standardized Clinical Assessment for Practitioners: A Primer. How standard scores, percentile ranks and the standard error of measurement are meant to be read together.
Pearson (2008). WAIS-IV Score Report sample. What a real report contains: every composite paired with a percentile rank, a 95% confidence interval and a qualitative description, never a bare number.
Crawford, J.R., Garthwaite, P.H. & Slick, D.J. (2009). On percentile norms in neuropsychology. The Clinical Neuropsychologist, 23(7), 1173-1195. Three definitions of a percentile coexist in practice and can return different results for the same score.
Voncken, L., Albers, C.J. & Timmerman, M.E. (2019). Improving confidence intervals for normed test scores. Behavior Research Methods. Open access. Documents the mean 100, SD 15 metric and the uncertainty that norming from samples adds to any score.
Pearson (2024). WAIS-5, Wechsler Adult Intelligence Scale, Fifth Edition. The current adult battery, covering ages 16:0 to 90:11 across five cognitive domains.
Take the assessment
You get a profile, not a number
ACIS measures six CHC domains across 20 subtests and reports each one with its own normed score and confidence interval, so you can see where you are strong and where you are not.