Skip to main content

Pre-employment assessment: what actually predicts performance

EH
Expert Hire Team
September 16, 2026
Pre-employment assessment: what actually predicts performance
Share this article

A pre-employment assessment is any structured test a company runs before an offer to predict how well someone will do the job: aptitude tests, personality questionnaires, skills tests, work samples, or a structured interview. Here's the position this guide takes. Most of them measure a proxy for the work. The strongest one measures the work itself.

That distinction is the whole game, and most buyer's guides skip it. They hand you a catalog of test types and leave you to guess which one predicts anything. This guide leads with the evidence, then says where each type earns its place and where it just adds cost.

Key Takeaways

  • Most pre-employment assessments measure a proxy for the job. The strongest single method is a work sample, a test that has the candidate do a slice of the actual work.

  • Predictive validity is the only score that matters. In the Schmidt and Hunter meta-analysis, work samples and structured interviews sit at the top and generic personality tests near the bottom.

  • Personality and aptitude tests are not useless, but sold as the default they are the weakest common choice, and cognitive tests carry the heaviest adverse-impact risk.

  • Fairness is not a separate box. A test that predicts performance and produces a defensible record is the one that survives an EEOC challenge, a NYC Local Law 144 audit, or an EU AI Act review.

  • When you buy a tool, check for reasoning per criterion, an audit trail, ATS write-back, and a human in the decision, not a bare pass or fail.

What a pre-employment assessment actually is

A pre-employment assessment is a standardized way to measure a candidate before you hire them, run the same way for everyone so the results compare. When every applicant faces the same task and the same scoring, you rank people on one scale instead of on how well each interviewer happened to click with them.

The labels move around. You will see the practice called pre-employment testing, and a single instrument called an employment assessment test. The vocabulary varies by vendor; the question underneath it does not.

The trap is treating "standardized" and "predictive" as the same thing. They are not. A test can be perfectly consistent and still measure something with almost nothing to do with the job, and a rigid one screens out people who could have done it well. Harvard Business School's Hidden Workers report estimated that automated screening pushes millions of qualified candidates out of the running.

So the real question is narrower than "is it consistent." It is: does the score predict who will actually do the job well?

The main types of pre-employment assessment, and what each measures

The category splits into a handful of families. Here is what each one measures, plainly, before we rank them.

Cognitive and aptitude tests

These measure general reasoning, numerical or verbal ability, and problem-solving speed. They are quick to run and predict performance reasonably well on their own. The catch is that they tend to produce the largest score gaps across demographic groups, which is where adverse-impact risk concentrates. A high score tells you someone can reason, not that they can do this job.

Personality questionnaires

These score traits like conscientiousness and emotional stability, usually through self-report. They are popular because they are cheap to bulk-administer. On their own they are weak predictors, and self-report is gameable: candidates who guess what "good" looks like answer accordingly. Our fuller take on where these fit lives in the guide to psychometric testing for hiring.

Skills and work-sample tests

These have the candidate do a piece of the actual work: write and run code, debug a broken function, build a model, draft the email. Because the test is a slice of the job, the score maps directly to the thing you are hiring for. This is the family that predicts best.

Integrity tests

These predict counterproductive behavior such as theft and rule-breaking. They are more useful than their reputation suggests and pair well with a skills measure, but they answer a narrow question and are not a general fit test.

Structured interviews

A structured interview asks every candidate the same job-relevant questions and scores answers against a fixed rubric. Structure is the active ingredient. The same conversation run loosely predicts far less. We break down the mechanics in the structured interview software guide.

The only score that matters: does it predict performance

Every assessment claims to work. The way you check is predictive validity, a number between 0 and 1 that says how strongly the test score tracks later job performance. Higher is better. It comes from decades of research correlating pre-hire scores with real outcomes.

The reference point is the Schmidt and Hunter meta-analysis of selection methods, which pooled roughly 85 years of validity studies. The ranking is the useful part:

  • Work sample tests: about .54

  • General mental ability (cognitive) tests: about .51

  • Structured interviews: about .51

  • Integrity tests: about .41

  • Unstructured interviews: about .38

  • Conscientiousness (personality): about .31

  • Reference checks: about .26

  • Years of job experience: about .18

  • Years of education: about .10

Read the top and the bottom together. Work samples and structured interviews lead. Generic personality self-reports sit near the floor, barely ahead of counting years on a resume. Cognitive tests predict well on paper, but they carry the adverse-impact problem below, and still do not show the person doing the work.

That is the evidence the personality-and-aptitude incumbents tend to gloss over. If you are going to run a test, run the one that predicts best.

Where personality and aptitude tests help, and where they mislead

None of this makes personality and aptitude tests worthless. Used well, they add signal. Conscientiousness has real predictive value as one input among several, and a cognitive measure can be a fair tiebreaker for high-reasoning roles. The mistake is making them the whole assessment.

Three failure modes recur. Self-report is coachable, so a personality score often measures how well a candidate reads the test, not who they are. A generic aptitude battery is role-blind: the same reasoning test gets sold for a sales rep and a backend engineer. And a cognitive score with a big group gap becomes legal exposure the moment it drives a rejection you cannot defend on job-relatedness.

The honest read: use personality and aptitude as supporting inputs if you like, but do not let them stand in for watching someone do the work. A proxy is a fine tiebreaker. It is a poor foundation.

Why structured work samples and interviews predict best

Work samples win for a boring, powerful reason. The test is the job. When a candidate writes code that has to compile and pass real cases, or works a real ticket end to end, you are not inferring performance from a trait. You are watching a small version of it.

That is why the work sample sits at the top of the validity table, and it is the whole idea behind good coding assessment software for technical roles.

Structure is what makes it repeatable. A work sample scored by gut feel loses most of its edge. The gain comes from every candidate getting the same task and the same rubric. Structured interviews land near work samples for the same reason.

The strongest first round combines the two. Give the candidate a real task, then have a structured conversation about the choices they made. You get the doing and the reasoning on one scale. For that to hold up later the scoring has to be legible, which is why how the scoring and reasoning work matters as much as the task.

Fairness, adverse impact, and the record you can defend

Fairness is not a compliance chore you bolt on after picking a test. The test you pick is the compliance posture. Regulators do not ask whether your assessment felt fair. They ask whether it is job-related and whether it screens out protected groups at a disproportionate rate.

In the United States, the EEOC's guidance on employment tests and selection procedures sets the baseline through the Uniform Guidelines. The common yardstick is the four-fifths rule: if one group's selection rate falls below 80% of the highest group's rate, that is a signal of adverse impact you have to justify with evidence that the test predicts the job.

This is exactly where a generic cognitive battery gets a company into trouble, and where a job-relevant work sample is easier to defend.

The regulatory layer has grown past the EEOC. NYC Local Law 144 (LL144) requires a bias audit of automated employment decision tools plus candidate notice, and it is now a template other jurisdictions are copying. The EU AI Act classifies most hiring tools as high-risk and expects documentation of how the system works and how it is monitored. Illinois regulates AI use in video interviews through its consent law.

The practical takeaway is simple. Favor an assessment that is job-related on its face and that produces a written record of why each score was given. That record turns a bias audit or a candidate challenge from a scramble into a document you already have.

Our guide to reducing bias in hiring goes deeper on the audit posture. None of this is legal advice; run your program past counsel.

How to run a work sample at first round without booking a human every time

The usual objection to work samples is cost. A senior engineer sitting through a live exercise with every applicant does not scale, so teams fall back to cheap-but-weak proxy tests. That trade-off created the personality-test default in the first place.

You do not have to accept it anymore. Keep the work sample and lose the cost by automating the first round while keeping a human on the decision. In practice that is a structured, role-tuned interview with a live task, scored against a rubric, with a person reviewing before anyone is advanced or rejected.

This is where Expert Hire sits. Expert Hire is not a personality-test or cognitive-aptitude-battery vendor. It runs a structured, conversational AI interview tuned to the role, with a live coding exercise that executes the candidate's real code, scored against a transparent rubric, and reviewed by a human before any decision.

It is a work-sample-and-structured-interview instrument, at the predictive end of the table above, not another proxy test. It does not offer personality or aptitude scoring.

The point is not the tool. It is that the scaling excuse for weak tests has expired. You can give every candidate a real task at first round and still keep an engineer out of the loop until the final few.

What to look for when buying pre-employment assessment tools

Once you have decided to assess the actual work, the shopping list for pre-employment assessment tools gets short and specific. Judge a tool against these before you judge it on price.

  • Reasoning, not a verdict. The output should show why each score was given, with the evidence behind it, not a bare pass, fail, or percentile. A number you cannot explain is a number you cannot defend.

  • A human in the decision. The tool should support review before any advance or reject. A fully automated reject is the exact thing LL144 and the EU AI Act are built to scrutinize.

  • An audit trail. You want a durable record of the rubric, the score, and the rationale, exportable for a bias audit or a candidate request.

  • Job-relatedness on its face. The task should look like the work. If you cannot explain to a candidate why the test maps to the role, a regulator will have the same question.

  • ATS write-back. The result should land in your existing candidate record automatically, so the assessment is part of the pipeline and not a spreadsheet someone re-keys.

  • Candidate experience. A test that feels like the job respects the candidate's time and protects your employer brand. A test that feels like a hazing ritual costs you offers.

Run any tool, ours included, against that list. The grounding for how we approach scoring and validity is on our methodology page, written for the buyer who reads the evidence before the marketing.

Frequently asked questions

What is a pre-employment assessment? It is a standardized test given to candidates before an offer to predict how well they will do the job. It can measure aptitude, personality, integrity, or job skills, or it can be a structured interview. The useful ones predict on-the-job performance. The rest just add a step.

What are the main types of pre-employment tests? Cognitive and aptitude tests, personality questionnaires, skills and work-sample tests, integrity tests, and structured interviews. They differ sharply in how well they predict performance. Work samples and structured interviews lead the evidence; generic personality self-reports sit near the bottom.

Which pre-employment assessment predicts job performance best? Work-sample tests, closely followed by structured interviews, according to the Schmidt and Hunter meta-analysis. Both work because they measure job-relevant behavior directly instead of inferring it from a trait or a generic reasoning score.

Are pre-employment tests legal? Yes, when they are job-related and do not cause unjustified adverse impact. The EEOC's Uniform Guidelines set the US baseline through the four-fifths rule, and newer rules such as NYC Local Law 144 and the EU AI Act add audit and documentation duties for automated tools. Keep a defensible record and consult counsel for your jurisdiction.

Do personality tests work for hiring? As a supporting input, sometimes. As the main assessment, weakly. Traits like conscientiousness carry some predictive value, but self-report is coachable and the validity is low next to a work sample. Use personality data as a tiebreaker, not a foundation.

How can we run work samples without a huge time cost? Automate the first round and keep a human on the decision. A structured, role-tuned interview with a live task, scored on a rubric and reviewed by a person, gives every candidate a real work sample without booking a senior reviewer for each one.

The assessment worth running is the one that measures the work

If you take one thing from this guide, take the ranking. Most pre-employment assessments measure a proxy, and the strongest single method is a work sample that has the candidate do a slice of the actual job. Personality and aptitude tests can support that decision. They should not be the decision.

So assess the work, keep it structured, keep a human on the call, and keep a record you could hand to an auditor without flinching. If you want to see what that looks like in practice, open a sample scorecard on the AI interview platform and read the rubric, the transcript, and the reasoning behind each score before you decide whether it holds your bar.

Ready to Transform Your Hiring?

Start your free trial to see how Expert Hire can help you screen candidates faster and smarter.

Share this article