
The free 15-minute test shows your IQ band and strongest cognitive domain — the numbers this article keeps referring to.
The email arrives on a Tuesday and gives Vanessa forty-eight hours. Her application for a financial analyst role has cleared the resume screen. The next step is a fifteen-minute test she has never heard of. She spends the evening hunting for practice questions. What she never finds is the thing that would have helped her: a description of the one-page report her score becomes, who opens it, and what number they compare it against.
A pre-employment cognitive test is not scored against a passing grade. It is scored against a job. The vendor turns your answers into a raw score and a percentile. Your prospective employer decides, before you ever sit down, what that number has to clear for the role you applied to. The same score can advance you at one company and end your application at another on the same afternoon.
Almost everything written for candidates is about the questions. This guide is about the other side of the transaction: which test you are facing, what the report on the hiring manager's screen contains, how a number becomes a yes or a no, what the law lets you ask for, and what happens to the result afterwards.
Here are the figures that govern the process, each with the source that published it. Every one is checkable, which is more than can be said for most of what circulates about employer testing.
| Figure | Source | As of | |
|---|---|---|---|
| CCAT format | 50 questions, 15-minute limit; 'very few people finish all 50 items' | Criteria Corp, CCAT Score Report Guide | 2026 |
| CCAT benchmark | Raw score 24 = 50th percentile against the global norming group | Criteria Corp, sample score report | Jan 2026 |
| Job-family ranges | Suggested raw scores from 17–50 (Retail Sales) to 24–50 (Legal, Data Analytics) | Criteria Corp, over 700,000 CCAT assessments | 2026 |
| PI Cognitive scale | Raw score converted to a 100–450 scaled score, compared to a job Cognitive Target | Predictive Index documentation | 2026 |
| Validity today | Mean corrected validity of .22 for predicting job performance | Sackett et al., Journal of Applied Psychology | 2024 |
| Validity before | .51, the figure that drove 25 years of hiring practice | Schmidt & Hunter, Psychological Bulletin | 1998 |
| Prevalence | 56% of 1,688 surveyed SHRM members used pre-employment assessments of any type | SHRM skills-based hiring survey | Aug 2022 |
| Adverse impact | A selection rate below 4/5ths of the top group's rate is evidence, but not a legal definition | EEOC, 29 CFR 1607.4(D) | 1978 |
Two of those rows belong together. The validity of a cognitive test fell sharply over the last four years, and the tests are still in wide use. That is not a contradiction. Understanding why is the most useful thing you can know walking into one.

For twenty-five years the number that justified cognitive testing was .51. That was the correlation between general mental ability and job performance reported by Schmidt and Hunter (1998), and it built the industry.
It has since been revised. Sackett and colleagues showed that the range-restriction correction behind the classic estimates had overcorrected. Repairing it dropped the validity to roughly .31 (Sackett et al., 2022). A follow-up drew on contemporary samples alone and reported a mean corrected validity of .22 (Sackett et al., 2024).
So why is the test still on the table? The answer is comparative. The same correction was applied to every selection method, and the whole field compressed together rather than cognitive ability dropping out of it. The practical advice that followed was to combine predictors instead of leaning on one (Sackett et al., 2023).
There is also a much less scientific reason. A cognitive test is cheap, fast, and scales to thousands of applicants without a human reading anything. For a role drawing several hundred applications, it is the only stage whose cost per candidate is close to zero. Critical reviews of testing in graduate hiring make the same point. The volume problem, not the validity evidence, is what keeps these tests in the funnel (Woods et al., 2023).
The evidence has real defenders too. A 2025 review argued that even the revised estimates leave cognitive ability among the better-supported predictors available, and that dropping it would trade a measurable signal for an unmeasurable one (Kulikowski, 2025).
Four assessments cover most corporate cognitive screening in the US and UK. What follows is what each one is for from the employer's side. The tactics, meaning pacing and question types and what to drill, belong to our breakdown of employer cognitive tests, which goes deeper on each than this guide will.
| Format | What the employer gets | Typically used for | |
|---|---|---|---|
| CCAT (Criteria Corp) | 50 questions / 15 min | Raw score, percentile, three subscores, job-family range | Analytical, technical and management roles |
| Wonderlic | 50 questions / 12 min | Score plus a job-fit judgement against the role | High-volume and operational hiring |
| PI Cognitive Assessment | 50 questions / 12 min | 100–450 scaled score against a job Cognitive Target | Roles hired through the PI platform |
| SHL Verify G+ | Sectioned, adaptive | Percentile against a chosen comparison group | Graduate schemes, banking, large employers |
The CCAT measures what Criteria calls the ability to "solve problems, digest and apply information, learn new skills, and think critically," split across spatial reasoning, verbal ability, and math and logic. Wonderlic is the oldest of the four and leans more on verbal and numerical knowledge than on abstract reasoning. The PI Cognitive Assessment is the most time-pressured, at 12 minutes for 50 items. SHL Verify G+ is sectioned and adaptive rather than one sprint, and it reports percentiles against a comparison group the employer picks. That is why an SHL result is much harder to read out of context than a CCAT raw score.
One structural fact matters more than any of the differences. These tests correlate strongly with each other because they all estimate the same underlying thing. Preparing for the format of one gives you most of what preparing for another would. Preparing for the content gives you almost nothing, because there is no content to learn.
This is the part no prep site publishes, and it sits in the vendor's own documentation.

The CCAT score report opens with your name, the position you applied for, the date you sat the test, and a Test Event ID. Below that sits the Results Summary. It gives a Raw Score, which is the number of questions you answered correctly, and a Percentile ranking against Criteria's global norming group.
In the sample report Criteria publishes, a raw score of 24 corresponds to the 50th percentile. Answering fewer than half the questions correctly puts a candidate at the median.
The next block splits your performance into three separate percentile rankings: spatial reasoning, verbal ability, and math and logic. In Criteria's sample, a candidate at the 50th percentile overall sits at the 58th for spatial, the 62nd for verbal, and the 24th for math and logic. A hiring manager filling a quantitative role can see that profile. You will almost certainly never see it.
The Predictive Index works differently in a way worth knowing. Your raw score becomes a scaled score between 100 and 450. PI's own guidance is that the scaled score, not the number of items you attempted, is what should inform decisions. The employer sets a Cognitive Target for the role through a separate job assessment, and your scaled score is compared against it. PI also tells administrators to communicate job fit or ranking rather than raw scores, which is a large part of why so few candidates ever learn their number.
There are three common decision rules. Which one an employer uses changes what your score needs to be.
The vendor benchmarks are unusually open for the CCAT. Criteria publishes suggested raw-score ranges by job family, calculated from over 700,000 assessments. The floor varies more than most candidates expect. Retail Sales, Customer Service Representative, Sales Representative and Service Technician all start at 17. Store Manager and Plant Operator start at 18. Management and Software Development start at 20. Project Manager and IT Systems start at 21. Engineering, Science and Senior Leadership start at 22. Legal and Data Analytics start at 24. Every range runs to 50 at the top. There is no such thing as scoring too well.
The gap between a floor of 17 and a floor of 24 is the whole practical spread of this test. It is decided by which job family the recruiter tagged your application with. That is worth sitting with. It is also why comparing your score to a friend's, or to a number on a forum, tells you almost nothing.
Professional standards treat a threshold as a technical and policy decision rather than a convenient percentile. The Standards for Educational and Psychological Testing (2014) require that the reasoning behind a cut score be documented, and that the standard error of measurement be reported near it. The reason is simple: a score just below a threshold may not mean anything different from one just above it. SIOP's Principles for the Validation and Use of Personnel Selection Procedures (5th edition, 2018) takes the same view and prescribes no universal number. The wider selection literature agrees. How a score is used drives validity and adverse impact at least as much as which test was picked (Van Iddekinge et al., 2023).
Yes, but far less than the prep industry sells, and the gains come from one specific place.
Practice moves scores mainly by removing the novelty penalty. You learn the question types, you build a pacing reflex, and you stop losing thirty seconds to the interface. Meta-analytic work on retest effects finds them on both kinds of task, and finds them running out: working memory tests gain a weighted average of g = 0.28 from a first sitting to a second (Scharfen, Jansen, & Holling, 2018), while across 174 samples of cognitive ability tests the gains are significant but stop accumulating after the third administration (Scharfen, Peters, & Holling, 2018). That is the signature of familiarity, not of new ability.
What preparation does not do is raise the ability the test is estimating. The full evidence review sits in our guide to how much you can improve a cognitive test score, which covers what works, what the marketing claims that the research does not support, and a four-week timeline. This guide will not re-argue it.
What this guide adds is the employer-side limit that the coachability research does not cover. There is a ceiling on how much practice is safe. Familiarity with the format helps you. Familiarity with the live test triggers the Invalid Result flag described above, and a report reaching a recruiter with a validity warning on it is worse than a mediocre score.

If a disability affects your ability to take a timed test on equal terms, the Americans with Disabilities Act entitles you to a reasonable accommodation in the application process. That includes tests. The EEOC's guidance for candidates, Job Applicants and the ADA (EEOC-NVTA-2003-4), sets out how it works, and most of it surprises people.
You do not have to use the phrase "reasonable accommodation." You only have to tell the employer that you need a change to the process because of a medical condition. The request may be spoken or written. Someone else can make it for you, such as a family member, a health professional or a job coach.
Timing is what matters most. Ask as soon as you know a test is coming, because extra time, alternative formats and assistive technology all take arrangement.
Where your disability and your need are not obvious, the employer may ask for reasonable documentation explaining the disability and why the accommodation is needed. That is the limit. They may not probe the nature or severity of the condition beyond establishing the need. There is no requirement to hand over a diagnostic label or your medical records, and anything you do disclose must be kept confidential.
The employer's matching duty appears in the EEOC's Employment Tests and Selection Procedures (EEOC-NVTA-2007-2, December 2007). Tests must be given in a format that does not require the use of an impaired skill, unless the test is designed to measure that skill. That exception is narrower than employers sometimes claim. Processing speed under time pressure is arguably what a fifteen-minute test measures. The reading, the mouse control and the interface are not, and those can be accommodated. In one EEOC settlement, DaimlerChrysler provided readers and audiotape versions of tests to applicants with learning disabilities.
Employment testing sits inside a body of law that exists because tests have been used badly.
Griggs v. Duke Power Co., 401 U.S. 424 (1971), established that a practice neutral on its face but discriminatory in effect breaks Title VII unless it bears a manifest relationship to the employment in question. The Court's formulation is still the standard. The touchstone is business necessity, and a practice that excludes a protected group without being shown to relate to job performance is prohibited. Albemarle Paper Co. v. Moody, 422 U.S. 405 (1975), applied that rule to tests. A test with a discriminatory effect must be shown, by professionally acceptable methods, to predict important elements of the work.
The Uniform Guidelines on Employee Selection Procedures (29 CFR Part 1607, 1978) put this into practice. The much-cited four-fifths rule at § 1607.4(D) says a selection rate below 80% of the top group's rate will generally be regarded by enforcement agencies as evidence of adverse impact. The EEOC's own interpretation is emphatic that this is a rule of thumb and not a legal definition. Smaller differences can count as adverse impact, and larger ones may not.
The Guidelines also place the validation burden on the employer, not the vendor. As the EEOC's 2007 guidance puts it, vendor documentation may help, but "the employer is still responsible for ensuring that its tests are valid." In the Ford ATSS matter, a validated test kept producing a statistically significant disparate impact where less discriminatory alternatives existed. Ford settled for $8.55 million.
Two further points bear on your position. Group differences on cognitive tests are real and well documented, and they interact with the conditions under which the test is taken. The stereotype threat literature, reviewed across two decades of studies, describes mechanisms that operate at the moment of testing rather than in underlying ability (Pennington et al., 2016). Separately, in New York City, Local Law 144 requires employers using an automated employment decision tool to commission an annual independent bias audit, publish a summary of it, and notify candidates at least 10 business days before the tool is used. If you applied to a New York City role and no such notice arrived, that is information about the employer.

Three things about the aftermath catch candidates out.
You will probably never be told your score. Wonderlic's candidate material states that a Candidate Feedback Report is the sole result available to applicants, and that numerical scores are not provided unless the employer shares them. PI advises its administrators to keep actual scores confidential and to communicate fit or ranking instead. The asymmetry is total. They get three subscores and a job comparison. You get an email.
A retake is the employer's decision, not your right. No major publisher grants candidates a general retake entitlement. Wonderlic tells employers that retesting is warranted where a result carries a retest warning, and that testing a candidate more than twice on a component is inadvisable. If you want another attempt, you are asking a recruiter, not a vendor.
Your score does not travel with you. A result sits in the account of the employer that administered it. There is no portable cognitive credential and no verified cross-employer expiry rule. A new company will simply test you again. That cuts both ways, because a bad result at one firm does not follow you to the next.
The one thing that does travel is what the process did to you. Applicant reactions to selection procedures are a well-studied outcome in their own right, and they shape whether candidates stay with an employer at all (Woods et al., 2019). A process that felt arbitrary and told you nothing is a data point about that organisation, and you are allowed to weigh it.
One last thing does travel, in the wrong direction. Unproctored testing has an AI problem, and the employer answer is not better detection but verification testing: a shorter proctored re-test later in the process, benchmarked against your first result. Large language models now score well enough on ability items to push employers toward that step (Hickman et al., 2024), so using AI on the unproctored test mostly buys you an unexplainable gap between two scores. Our guide to cognitive tests at McKinsey, BCG and FAANG covers how the elite firms build that sequence.
A corrected validity of .22 means cognitive ability accounts for a modest share of the variance in job performance. The rest sits in things a fifteen-minute test does not try to measure.
This is not consolation. It is the argument the selection research makes. No single predictor carries a hiring decision, and a system built on one is a worse system (Sackett et al., 2023). If an employer treats a fifteen-minute score as a verdict rather than a filter, they are using their own instrument wrongly. That tells you something about how they will use other measurements once you are inside.
For a sense of where cognitive demand concentrates by occupation, our cognitive demand index for 100 jobs maps the evidence role by role. Our guide to whether you are smart enough for elite careers takes on the threshold question head on.

Given all of the above, a proportionate plan is short. Take one timed diagnostic under real conditions, at a desk, no calculator, no pauses, to find where the format costs you time. Do two or three more timed runs on practice material, never on the employer's live test. Decide your skip threshold in advance, because the most expensive habit is grinding on one hard item while three easy ones expire. Then sleep properly and take it in the morning, on a desktop.
That is a few days, not a month, and it captures nearly all of the available gain. Anything beyond it is buying reassurance. The detail sits in our preparation timeline guide and our test-day checklist.
Knowing your own profile is worth more than extra drilling. Criteria splits its report into spatial, verbal, and math and logic for a reason. Those three seldom sit at the same level in one person. Knowing that math and logic is your slow domain changes which questions you skip in the first four minutes, and that decision is worth more points than another practice test. Our assessment measures five domains and reports them separately: fluid reasoning, verbal reasoning, quantitative reasoning, working memory and executive function. That is the same shape of information the employer's report will hold about you.
Fifteen minutes, free to take, no sign-up to start. Your band and strongest domain stay free.
The article explains the idea; these measure it.