Aptitude usually refers to readiness or potential in a particular area, while intelligence refers to broader capacities shared across many cognitive tasks. In practice, test names are not perfectly consistent, so content, validation and intended use matter more than the label.
Aptitude tests target readiness for particular learning or work demands; intelligence batteries seek broader cognitive structure.
0 The Short Answer
Aptitude and intelligence are not two different substances inside a person's head. They are two different jobs a test can be built to do. An aptitude test is engineered to forecast one named outcome in one named domain: will this recruit finish electronics training, will this apprentice keep pace on a shop floor, will this student survive the first year of an engineering program. An intelligence battery is engineered for a different task: describe where a person stands across broad cognitive domains and estimate the general factor running through all of them. The raw material overlaps almost completely. The purpose, the score report and the reference population do not.
That reframing dissolves most of the confusion. People ask whether aptitude is "innate potential" and intelligence is "developed ability", or the reverse, and neither reading survives contact with actual test manuals. A mechanical comprehension test asks about pulleys and gear trains, content nobody is born knowing. A fluid reasoning subtest asks you to complete novel patterns you have never studied. Both are samples of current performance under standard conditions. What separates them is the question the publisher promised to answer, and the group your score gets compared against.
There is a second answer that the marketing copy for aptitude products rarely mentions, and it is the most interesting thing on this page. When researchers checked whether differentiated aptitude profiles predict real outcomes better than a single estimate of general cognitive ability, the extra profile detail earned far less than anyone expected. That finding has been replicated in military training pipelines, in pilot selection and in civilian employment research since the early 1990s, and it changes how you should read any product promising to map your unique aptitude fingerprint.
What this page is and is not
This is an explanation of how two families of instruments are built, normed and used in hiring and career guidance. It is not career advice, not a hiring recommendation, and not a substitute for a supervised professional assessment. ACIS is an online self assessment that reports a domain profile with confidence intervals; it is not a clinical or diagnostic instrument.
1 What "Aptitude" Means in a Test Manual
Strip away the everyday usage, where aptitude means roughly "knack for", and the technical definition is narrow and practical. An aptitude is the capacity to acquire competence in a defined domain, and an aptitude test is an instrument validated to predict a specific criterion: grades in a training course, supervisor ratings after six months, time to reach journeyman standard, pass rates in a licensing exam. The criterion is not decorative. It is what the test is for, and without it the word aptitude has no technical content at all.
This is why aptitude testing is fundamentally a prediction business rather than a description business. Ask a psychometrician whether a mechanical aptitude test is any good and the answer will not be a philosophical account of mechanical ability. It will be a correlation coefficient against a stated outcome in a stated population, with a stated sample size and a stated time gap between test and criterion. Change the criterion and the same test can go from excellent to useless without a single item changing.
Two consequences follow that trip people up constantly. First, an aptitude test is only meaningful for the outcome it was validated against. A clerical speed and accuracy test that predicts data entry throughput tells you nothing about whether someone will make a good supervisor, and the manual will say so if you read past the cover. Second, aptitude scores are usually reported against occupational or educational norm groups rather than the general population. A spatial score at the 50th percentile among applicants to a drafting program means something quite different from the 50th percentile among all adults, and confusing the two is the single most common misreading of aptitude results.
The founders of the field were explicit about this. George Bennett, Harold Seashore and Alexander Wesman, the authors of the Differential Aptitude Tests, published a piece in Occupations: The Vocational Guidance Journal in 1952 with the blunt title "Aptitude testing: does it prove out in counseling practice?" They were not selling certainty. They were asking, in print, whether their own product delivered on its predictive promise. That is the correct posture, and it has been largely abandoned by the modern consumer aptitude quiz.
2 Where Aptitude Testing Came From, and Why It Looks the Way It Does
The shape of aptitude testing is not the product of a theory of mind. It is the product of two large institutional problems that had to be solved at scale, and the instruments still carry the fingerprints of both.
The first problem was military classification. Armed forces do not simply need to know who is bright. They need to sort hundreds of thousands of people into dozens of technical specialties in a matter of weeks, and the cost of a wrong assignment is a failed training course. That pressure produced batteries built around occupational clusters rather than around cognitive theory. The Armed Services Vocational Aptitude Battery, administered through the United States Military Entrance Processing Command, was first introduced in 1968 and adopted across all branches by 1976. Its nine subtests include General Science, Arithmetic Reasoning, Word Knowledge, Paragraph Comprehension, Mathematics Knowledge, Electronics Information, Automotive and Shop Information, Mechanical Comprehension and Assembling Objects. Four of them, combined as Arithmetic Reasoning, Mathematics Knowledge and a verbal composite drawn from Word Knowledge and Paragraph Comprehension, form the Armed Forces Qualification Test used for enlistment eligibility. The remaining subtests feed branch specific composites that gate entry into particular specialties.
Notice what that design implies. Subtests like Automotive and Shop Information and Electronics Information are frankly measures of prior exposure to a subject area. Nobody pretends otherwise. They sit in an aptitude battery because knowing something about engines predicts who finishes engine school, and prediction is the point.
The second problem was public employment placement. Government labor agencies had to refer millions of job seekers to employers across every occupation in the economy, which required a general purpose battery yielding many separate scores. The General Aptitude Test Battery was built for exactly that, and its use became contentious enough that the National Research Council published a full length review in 1989, "Fairness in Employment Testing: Validity Generalization, Minority Issues, and the General Aptitude Test Battery" through the National Academies Press. The report examined whether validity evidence gathered on some occupations could legitimately be generalized to others, and what the equity consequences of the referral system were. It remains the most serious public audit any aptitude battery has received.
3 How a Multiple Aptitude Battery Is Built and Reported
Open the score report from a multiple aptitude battery and you get a fan of separate scores, typically eight to twelve, each labeled with an occupational flavor: verbal reasoning, numerical ability, space relations, mechanical reasoning, clerical speed and accuracy, spelling, language usage, abstract patterns. The visual signature is a profile chart with peaks and valleys, and the peaks are the selling point. The implied promise is that your shape, not your height, tells you where to go.
The construction logic runs backward from jobs. Developers begin with an occupational map, ask what those jobs appear to demand, write subtests sampling those demands, then validate each subtest or composite against training and performance criteria in the relevant occupations. Composites are then assembled from whichever subtests predict best, which is why military line scores mix subtests in combinations that look arbitrary from a theory standpoint but are empirically tuned.
Three design features follow from that logic and matter for interpretation:
Content specificity is deliberate. Items are written to look like the work. Mechanical items show levers and gears, clerical items show name and number matching. Face validity is not vanity here; it keeps applicants engaged and keeps employers convinced the test is job related.
Speed often carries real weight. Clerical and checking subtests are frequently pure speed measures with easy items and tight limits, closer to processing speed than to reasoning. Scores on them move with practice and motivation more than most people assume.
The profile is the product. Marketing, counseling protocols and software all revolve around comparing your peaks to occupational requirement patterns. That is precisely the feature the research in section 6 puts under strain.
Worth adding: multiple aptitude batteries usually report against applicant or student norms and often use stanines or percentile bands rather than a 100 point scale. If you have only ever read IQ style index scores, an aptitude profile can look unfamiliar even when it is measuring overlapping ground.
4 How an Intelligence Battery Is Built and Reported
An intelligence battery inverts the logic. Instead of starting from a list of jobs, it starts from a structural model of cognitive ability, most commonly the Cattell Horn Carroll framework, and builds subtests to sample the broad domains that model specifies: verbal comprehension, fluid reasoning, visual spatial processing, working memory, processing speed and quantitative reasoning. Subtests are grouped into indexes, indexes are combined into a composite estimate of general ability, and the whole structure is checked with factor analysis to confirm that the subtests cluster the way the model says they should. That composite is a deviation score, read against the distribution for people of the same age, and not the ratio of mental age to chronological age that early scales printed; the mental age convention collapsed for adults because cognitive development is not proportional, so no single mature mental age keeps pace with birthdays.
The reference group is the general population, stratified by age and often by other demographic variables, which is why the score scale has a fixed mean of 100 and a standard deviation of 15. That scale is doing real work: it turns any subtest, whatever its content, into a statement about where you stand among adults in general. A score of 130 corresponds to roughly the top 2 percent, about 1 person in 44; 145 falls near 1 in 740; 160 is near 1 in 31,000. Those rarity conversions only hold because the norm group is the population rather than a self selected applicant pool.
The second structural difference is the confidence interval. Serious cognitive reports do not hand you a point estimate and walk away. They give each index a band, because a single administration measures with error, and a difference of four points between two indexes is usually noise rather than a meaningful peak. Aptitude profile charts frequently invite exactly the peak reading that confidence intervals are designed to discourage.
The third difference is what gets emphasized. Intelligence batteries report both the profile and the general factor, and the general factor is usually the most reliable number on the page, because it pools variance across every subtest. Aptitude batteries often bury or omit the general composite entirely, since their commercial story is about differentiation. Both instruments contain the same general factor. Only one of them shows it to you clearly.
5 Aptitude, Intelligence and Achievement Side by Side
The third term in this family is achievement, and leaving it out is why the aptitude versus intelligence question stays muddy. A reading comprehension passage can appear in all three kinds of test with barely a word changed. What differs is the contract the publisher signs. A fourth label has since been layered on top of the three, and it sorts by breadth rather than by promise: the cognitive ability test is defined as a broad estimate of general mental ability across several domains, against which an aptitude test is the narrow instrument aimed at one ability at a time.
Dimension
Aptitude battery
Intelligence battery
Achievement test
Question it answers
Will this person succeed in a specific course, trade or role?
Where does this person stand across broad cognitive domains?
What has this person already learned in a defined curriculum?
Time direction
Forward looking, tied to a future criterion
Present standing, no single criterion promised
Backward looking, tied to instruction already delivered
Typical norm group
Applicants, trainees or students in that field
General population, stratified by age
Grade or course cohort
Score format
Profile of eight to twelve scores plus occupational composites
Domain indexes with confidence intervals plus a general composite
Mastery levels, scaled scores or grade equivalents
Evidence that matters most
Criterion related prediction against the stated outcome
Factor structure, reliability, population representativeness
Curriculum coverage and content alignment
Main failure mode
Profile peaks read as meaningful when they are measurement noise
Composite treated as a fixed personal trait rather than a dated estimate
Mistaken for a measure of capacity rather than of exposure
Read down the last row and you have the practical difference between the three. Each family fails in its own characteristic way, and knowing which failure you are exposed to is more useful than memorizing definitions.
6 The Finding That Complicates Everything
Here is where the field gets genuinely uncomfortable, and it is the part most articles on this topic skip. If specific aptitudes are real and distinct, adding them to a prediction equation should improve that prediction beyond what general cognitive ability already delivers. Researchers with access to enormous military datasets tested exactly that, and the answer was not what practitioners hoped.
Malcolm Ree and James Earles published a run of studies through the early 1990s using large United States Air Force samples. In Personnel Psychology in 1991 they reported on predicting training success, and in Current Directions in Psychological Science in 1992 they summarized the programme under a title that leaves nothing to interpretation: "Intelligence Is the Best Predictor of Job Performance". With Mark Teachout in the Journal of Applied Psychology in 1994 they extended the analysis to job performance itself, and in the same journal that year Michele Olea and Ree applied the approach to pilot and navigator criteria. The recurring conclusion across those papers is captured by the phrase that became their shorthand: not much more than g. Specific aptitude scores, once general ability was accounted for, added small increments to prediction across a wide range of technical specialties.
Set that alongside the broader employment literature. Frank Schmidt and John Hunter's 1998 review in Psychological Bulletin, covering eighty five years of selection research, placed general mental ability at the centre of the prediction toolkit and evaluated other methods largely by what they added on top of it. That paper shaped hiring practice for two decades.
It also needs a modern correction, and the correction is instructive rather than dismissive. Paul Sackett, Charlene Zhang, Christopher Berry and Filip Lievens published a reassessment in the Journal of Applied Psychology in 2022 arguing that standard corrections for range restriction had been systematically overcorrecting. Their revised estimates came in roughly .10 to .20 lower than the earlier meta analytic figures, with structured interviews emerging as the top ranked procedure and cognitive tests placing lower than the 1998 summary implied. The rank ordering shifted; the underlying point that specific aptitude profiles struggle to earn their keep above general ability did not get overturned by it. Prediction in this domain is simply weaker across the board than the field believed.
7 Why the Increment Is So Small, and Where Specific Abilities Still Win
The result sounds paradoxical until you look at how the scores are correlated. Cognitive tests of almost any content correlate positively with each other, a pattern known as positive manifold. A mechanical comprehension score and a verbal reasoning score share a great deal of variance before either one gets near a criterion. When you enter general ability into a prediction equation first, it absorbs most of what the specific scores had to offer, and the leftover unique variance in each aptitude is small, less reliable than the composite, and frequently uncorrelated with the outcome. Employers absorbed that lesson commercially long before most career counsellors did, which is why the screeners candidates actually face in hiring are single-score instruments such as the Wonderlic (50 questions in 12 minutes) and the CCAT (50 in 15) rather than differentiated profiles.
Four further mechanics deserve naming, because together they explain why differentiated profiles keep underperforming:
Reliability of difference scores is poor. The gap between two index scores carries the measurement error of both. A profile peak needs to be considerably larger than most people assume before it is distinguishable from noise, which is why well constructed instruments print intervals around every index.
Range restriction bites hardest on specifics. Selection systems already screen on general ability, so within a group of accepted trainees the general factor varies less, and the residual specific variance has correspondingly little room to demonstrate anything.
Criteria are broad, predictors are narrow. Training grades and supervisor ratings pool everything a person does. Broad outcomes are predicted best by broad predictors, which is a statistical fact rather than a claim about human nature.
Practice and exposure move specific scores. Content flavored subtests reward prior familiarity, so part of a peak reflects your biography rather than a stable ability difference.
None of this means specific abilities are fictional. The clearest counterexample is spatial ability. Jonathan Wai, David Lubinski and Camilla Benbow published an analysis in the Journal of Educational Psychology in 2009 aligning more than fifty years of longitudinal data, and found spatial ability carried genuine predictive weight for entry into and achievement within science and engineering fields, weight that verbal and quantitative measures alone had been missing. Their point was partly institutional: spatial ability is rarely assessed in the selection systems that route people into those careers, so a real signal goes unused. That is the honest version of the specific ability case, and it is narrower and better evidenced than the version printed on career quiz landing pages.
8 Real Aptitude Batteries and What They Are Actually For
Abstractions get slippery, so here are the instruments themselves, with their institutional purpose attached.
The Armed Services Vocational Aptitude Battery. The largest aptitude testing operation in the world by volume, run through the United States Military Entrance Processing Command. It does two jobs at once: gatekeeping through the Armed Forces Qualification Test, and specialty assignment through branch composites. It is also offered in schools as a career exploration tool, which means a great many civilians have taken a military classification instrument without ever intending to enlist.
The General Aptitude Test Battery. Built for public employment services to refer job seekers across the full breadth of the labor market, and the subject of the National Research Council's 1989 review cited earlier. Its history is the clearest case study in what happens when an aptitude battery is used at national scale and its consequences get audited in public.
The Differential Aptitude Tests. Developed by Bennett, Seashore and Wesman, who were writing about the instrument in the vocational guidance literature as early as 1948. The DAT is the archetype of the school counseling battery: a profile handed to a teenager to inform course selection and career thinking rather than to gate a hire.
Mechanical comprehension tests. Bennett's mechanical comprehension work spawned a genre still used for maintenance, technician and skilled trade selection. Items present physical arrangements and ask what happens next, and the format has survived largely unchanged for decades because employers find it credible and it predicts trade training outcomes reasonably well.
The O*NET Ability Profiler. A useful cautionary entry. Published by the United States Department of Labor, Employment and Training Administration as a free career exploration instrument, it was retired on November 20, 2021, with technical assistance discontinued and materials kept only for archival and research use. Public aptitude infrastructure is not permanent, and the free tool a career guide recommends may no longer exist.
Across all five, note the pattern: every one is embedded in an institution with a decision to make. Aptitude testing outside such a context, sold direct to an individual with no criterion and no local validation, is a different product wearing the same vocabulary.
9 Aptitude Versus Achievement: The Distinction That Will Not Hold Still
The textbook line is that achievement tests measure what you have been taught and aptitude tests measure your capacity to learn. Hold that line against real items and it buckles within a minute. Word Knowledge on a military aptitude battery is vocabulary, which is learned. Reading comprehension appears on aptitude batteries, intelligence batteries and school achievement tests alike. Electronics Information is unambiguously a knowledge test sitting inside an aptitude battery without apology.
What actually separates them is a pair of practical choices. The first is the criterion: an achievement test is answerable to a curriculum that has already been delivered, while an aptitude test is answerable to an outcome that has not happened yet. The second is the norm population: achievement scores compare you to others who received the same instruction, aptitude scores compare you to others entering the same field, and intelligence scores compare you to adults in general. The same forty items can serve all three roles depending on which comparison group the publisher standardized against and which validity study the manual reports.
This is not sloppiness on the field's part. It is a consequence of the fact that every cognitive test measures developed ability at the moment of testing. There is no instrument anywhere that reads capacity uncontaminated by learning, and any product claiming to measure pure innate potential is making a claim no psychometrician would sign. The fluid and crystallized distinction is the closest the field comes, and even fluid reasoning tasks are sensitive to schooling, test familiarity and strategy.
The practical upshot for a reader is a reordering of questions. Instead of asking whether a test measures aptitude or achievement, ask what outcome it was validated against, who is in the norm group, and how long ago the norms were collected. Those three answers determine what your score means. The label on the box determines almost nothing.
10 Six Numbers Worth Carrying Around
Concrete figures travel better than definitions. These six anchor the landscape described above, and each one is attributable.
1968
Year the Armed Services Vocational Aptitude Battery was introduced. All branches had adopted it by 1976, making it the longest running large scale aptitude program in operation.
9 subtests, 4 for the gate
The ASVAB's nine subtests feed specialty composites, but only Arithmetic Reasoning, Mathematics Knowledge, Word Knowledge and Paragraph Comprehension determine enlistment eligibility.
1989
Publication year of the National Research Council's review of the General Aptitude Test Battery, the most thorough public examination of validity generalization in employment testing.
.10 to .20
Approximate reduction in mean validity estimates reported by Sackett, Zhang, Berry and Lievens in 2022 after correcting for systematic overcorrection of range restriction.
50 years of data
Span of longitudinal evidence Wai, Lubinski and Benbow aligned in 2009 to show spatial ability predicts science and engineering outcomes that other measures miss.
1 in 44
Rarity of a score of 130 on the population scale with mean 100 and standard deviation 15. Aptitude batteries rarely use this scale, which is why their percentiles are not comparable.
The last card is the one people misuse most often. A percentile from an applicant pool and a percentile from a general population sample are different quantities, and moving between them without adjustment inflates or deflates a result by an amount nobody can estimate after the fact.
11 What an Aptitude Profile Can Honestly Tell a Career Decision
Vocational guidance was the original consumer application of aptitude testing, and it is where the gap between what these instruments deliver and what people want from them is widest. A teenager holding a profile chart wants a verdict: this is your field. The instrument can support something considerably more modest, and being clear about the boundary makes the tool more useful rather than less.
What a profile can support, with care: identifying a genuinely large gap between two domains, on the order of a full standard deviation or more, that survives repeat measurement; flagging a floor problem, such as quantitative reasoning too weak to survive a calculus sequence, which is a real constraint worth planning around; and giving a structured vocabulary for a conversation that would otherwise run on vague self impression.
What a profile cannot support: ranking careers by fit from small peaks, predicting satisfaction, or standing in for interest. That last one is the biggest practical failure in career testing. Ability and interest correlate weakly, and no ability battery has ever measured what a person wants to spend a decade doing. The best organized guidance systems keep the two instruments separate and treat their disagreement as information rather than error.
There is also a base rate worth stating plainly. Most people do not have dramatic profiles. Positive manifold means the common outcome is a broadly level profile at some overall height, with modest wobbles that shift between administrations. The distinctive spiky shape that career quizzes love to display is less frequent than their marketing implies, and when it appears it deserves confirmation before anyone builds a decade of plans on it. If your interest is really in the overall height, that is a different measurement question, and one covered under how score ranges are interpreted.
12 Which Question Is Your Online Test Actually Answering?
Bring this back to the practical situation: you are about to spend an hour on an online assessment and you want to know what you will get. The useful move is to identify which of three questions the instrument is built to answer, because no instrument answers all three. A fourth question exists as well; critical thinking measures ask how you use ability rather than how much of it you have.
If the question is "will I pass this specific program". You want a validated aptitude measure tied to that program, which in practice means the instrument the institution itself uses. A consumer test has no validation study against a training pipeline it has never seen, and no honest one will claim otherwise. Employers with real stakes handle this through local validation studies, which is a different exercise than taking a quiz.
If the question is "how do I compare to adults generally, and how is my ability distributed across domains". That is what a population normed cognitive battery is built for. You get index scores across verbal comprehension, fluid reasoning, visual spatial processing, working memory, processing speed and quantitative reasoning, each with an interval, plus a general composite. This is the ground the ACIS battery covers, across twenty subtests mapped to six domains, reported as a profile with intervals rather than a bare number. It is a self assessment, not a clinical evaluation, and the report says so.
If the question is "what have I learned in a subject". Neither family helps much. You want an achievement measure aligned to the curriculum, and the closest useful cousin is a practice exam for the qualification itself.
One more filter, applicable to any of the three. Ask what the norm group is, when it was collected, and whether results come with intervals. An instrument that will not tell you who you are being compared against has not measured you against anything. If you want to see how a domain profile reads before committing time to one, the example report below shows the format, and the relationship between measured ability and work outcomes is treated separately in the page on cognitive ability and job performance.
No, but the overlap is larger than most people expect. The words mark different testing purposes rather than different mental substances, and scores from the two families correlate substantially.
Does an aptitude test measure innate talent?
No instrument does. Every cognitive measure samples performance as it stands on test day, shaped by schooling, practice and familiarity with the format. Claims about pure inborn capacity have no measurement behind them.
Which is better for choosing a university course?
Neither on its own. Prior grades in the relevant subject usually beat both, and an interest inventory adds information that no ability measure contains.
Why do employers still use aptitude batteries if general ability predicts almost as well?
Applicants and managers accept tests that look like the job, legal defensibility is easier when content resembles work tasks, and some batteries were adopted long before the research questioning differentiated profiles appeared.
Can I prepare for an aptitude test?
Practice reliably raises scores on speeded and knowledge flavored subtests, and familiarity with item formats helps everywhere. That is exactly why an untimed practice score should not be treated as your real standing.
What does "incremental validity" mean?
It is the extra predictive accuracy a new measure buys once you already have the existing ones in the equation. A measure can correlate with an outcome and still add nothing incrementally.
Is the ASVAB an IQ test?
It is not designed or normed as one, though its qualification composite correlates strongly with general ability measures. Its scores are referenced to military applicant populations, not to adults at large.
Are aptitude percentiles comparable to IQ percentiles?
Usually not. If the reference sample is applicants to a field rather than the general population, the same percentile describes a different position, and there is no reliable way to translate between them afterward.
Does a high mechanical score mean I should become an engineer?
It is one weak input among several. Exposure to tools and machines inflates such scores, and career fit depends on interest, working conditions and training access that no ability test samples.
Why do career quizzes produce such dramatic profiles?
Short subtests measure with wide error, and small differences look large on a chart with a compressed scale. A vivid shape is also more shareable, which shapes how these products get designed.
What is the difference between a general composite and a domain index?
A domain index summarizes a related cluster of subtests; the general composite pools across all of them. The composite is the more stable figure because pooling averages out subtest specific error.
Do aptitude tests predict job performance well?
Modestly, and less well than the field assumed before the 2022 range restriction reassessment. Prediction of any single person's outcome remains rough regardless of which instrument is used.
Is spatial ability worth measuring separately?
The longitudinal evidence for science and engineering paths is stronger than for most other specific abilities, and it is frequently absent from selection systems, which is why it is treated as a special case.
Can an aptitude test be biased?
Any instrument can produce group differences and can be used unfairly. Bias in the technical sense concerns whether prediction works equally well across groups, which is a separate empirical question from average score gaps.
What happened to the free government aptitude tool?
The federal career exploration profiler was withdrawn in late 2021 and is no longer supported. Guides written before then still recommend it, so check that any suggested instrument is still published.
How many domains should a serious cognitive battery cover?
Enough to sample the major broad abilities in a recognized structural model, with several subtests behind each so that index scores are stable. Coverage matters more than the raw count of subtests.
Should I trust a profile from a fifteen minute test?
For entertainment, fine. For anything you plan around, no. Short administrations yield index estimates too imprecise to support comparisons between domains.
Does age matter when comparing scores?
Considerably. Population norms are built by age band, and speeded tasks in particular shift across adulthood. A score without its age reference is uninterpretable.
Can training raise my underlying ability or only my score?
Improvements from practice tend to be specific to the trained task and transfer poorly to untrained ones. Expect a better score on that format rather than a broad shift.
Why do results differ between two tests I took?
Different norm samples, different content mixes, different reliability and ordinary day to day variation all contribute. Two results within their overlapping intervals are not actually in conflict.
What should I look for before taking any assessment seriously?
A stated reference population with a collection date, published reliability figures, intervals around every reported score, and a plain statement of what the instrument does not claim to do.
14 Sources Behind This Page
The claims on this page follow the published literature rather than our own assertions. These are the primary papers and reference bodies worth reading directly, with what each contributes.
Frey, M.C. & Detterman, D.K. (2004). Scholastic Assessment or g? Psychological Science, 15(6). The study linking SAT scores to measured cognitive ability, and the reason SAT conversions are estimates rather than measurements.
Kuncel, N.R., Hezlett, S.A. & Ones, D.S. (2004). Academic performance, career potential, creativity, and job performance: can one construct predict them all? Journal of Personality and Social Psychology, 86(1).
Nisbett, R.E. et al. (2012). Intelligence: new findings and theoretical developments. American Psychologist, 67(2). The broad APA review of what moves measured intelligence and what does not.
Plomin, R. & Deary, I.J. (2015). Genetics and intelligence differences: five special findings. Molecular Psychiatry. Open access. The standard modern review of what twin and DNA evidence does and does not show.
Gottfredson, L.S. (1997). Mainstream Science on Intelligence, the editorial signed by 52 researchers. Intelligence, 24(1), hosted by the University of Delaware. A consensus statement on what IQ tests measure.
Spearman, C. (1904). General intelligence, objectively determined and measured. American Journal of Psychology, full text at Classics in the History of Psychology. The paper where the g factor entered psychology.
Voncken, L., Albers, C.J. & Timmerman, M.E. (2019). Improving confidence intervals for normed test scores. Behavior Research Methods. Open access. Documents the mean 100, SD 15 metric and the uncertainty that norming from samples adds to any score.
Pearson Clinical Assessment Scientific Council (2023). Standardized Clinical Assessment for Practitioners: A Primer. How standard scores, percentile ranks and the standard error of measurement are meant to be read together.
Pearson (2008). WAIS-IV Score Report sample. What a real report contains: every composite paired with a percentile rank, a 95% confidence interval and a qualitative description, never a bare number.
Take the assessment
You get a profile, not a number
ACIS measures six CHC domains across 20 subtests and reports each one with its own normed score and confidence interval, so you can see where you are strong and where you are not.